From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "From DNA Design to DNA Slimming".
Jane: Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, we're looking at this paper today, "From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer," and it seems like they’ve tackled a really specific design problem in the field of genetic sequence optimization. Jane, can you give us the main idea from what we’ve read so far?
Jane: Absolutely, Tom. Basically, this research is focused on sequence slimming: finding an exact-length subsequence from an existing DNA strand while keeping its function intact. The core claim here is that they developed a method to do this using agentic discovery, which they call ERA, and it found a designer that performs better than random or greedy approaches for several transcription factor binding targets.
Lu: That's fascinating because the whole point seems to be moving beyond just substituting bases to actually figuring out which parts of the sequence can be removed without losing what makes it work. I wonder how they framed this as an exact-length, order-preserving subsequence problem, since that’s a very strict constraint for any designer to meet.
Meng: From an engineering standpoint, that constraint sounds really tough to implement reliably in a search algorithm. If you're only allowed deletions and must maintain exact length, the search space gets incredibly complex quickly. I'm curious how they handled that difficulty during the discovery process.
Lalam: I think what's really compelling is the way they framed it using an agentic program search, where ERA searches over executable designer programs, and each program searches over DNA subsequences. This nested structure sounds like a smart way to ensure that whatever they find is actually a valid sequence manipulation tool.
Tom: Exactly, Lalam. And the paper points out that this approach allows the output of the outer loop to be ordinary code that can be inspected and tested directly on biological tasks. It sets up a level of auditability that’s important for scientific discovery, you know?
Jane: And they defined their success using a metric called R, where R equals E(b) minus the embedding energy of the slimmed sequence relative to the original. They claim that an R value of one means retaining the source's predicted lift over background, and anything greater than one improves that effect.
Paper summary: Lu: The paper also mentions they used hard checks in addition to natural language instructions when constraining ERA to produce valid DNA slimmers. That suggests they had to be very careful about cheating the deletion-only requirement, which is a common hurdle in these types of sequence design tasks.
Meng: I see that hard constraint on the deletion-only aspect as a necessary safeguard when dealing with agentic systems; it stops them from accidentally trying substitutions or insertions, which would violate the core premise. From an engineering perspective, setting up those checks must have been quite rigorous.
Lalam: And after running that search, they discovered a program called GRADASLIM which uses a two-phase search strategy involving sampling deletions from low-salience positions in the first phase. That sounds like a very practical way to quickly narrow down the huge possibilities before moving into more precise refinement.
Tom: GRADASLIM is what they discovered, and it seems it splits the process into getting to the length constraint fast, and then doing an exact-length refinement phase. It’s a systematic way of tackling a very messy optimization problem.
Jane: And that refinement phase involves updates to a salience vector, which they describe not as a model gradient but as a reservoir of credit assignment based on performance in better-than-average candidates. It sounds like they’re using historical data from good candidates to guide the next deletion choice.
Lu: That concept of using salience as a form of credit assignment, diffusing small amounts of credit to neighboring positions to preserve motifs, seems very intuitive for maintaining local structure during a deletion-only search. It’s essentially teaching the system what parts of the sequence are most valuable.
Meng: I wonder how that salience vector is maintained throughout the entire process; keeping track of that information across thousands of iterations must require a lot of computational overhead to manage efficiently. I'm thinking about the practical resource demands when running something like this on real hardware.
Lalam: The paper shows that GRADASLIM evaluates five targets: E2F3, ELF4, MAX, MECOM, and RAD21 at lengths of four hundred base pairs and one hundred base pairs. That’s a solid set of benchmarks they used to test the designer's capabilities.
Tom: And the evaluation results are quite compelling; ERA achieved the "highest mean in nine/ten settings" and exceeded random deletion in all ten targets. It also outperformed greedy methods for every target at four hundred base pairs, with paired bootstrap intervals showing positive results for every comparison.
Paper summary: Jane: That is significant data; the study concludes that ERA’s slimmed sequences actually improve the predicted lift over background when compared to their full-length sources, with the mean retained effect exceeding one in every setting they tested.
Lu: That suggests that this technique isn't just finding shorter sequences; it’s actively optimizing the biological function, which is what we were hoping to see from these sequence modification methods. The ability to expose which sequence features drive predicted activity is a major piece of information here.
Meng: If this works consistently across different targets and lengths, that moves it closer to being a viable tool for actual use in designing regulatory elements, which is what we need for practical applications. I'm interested in the practical implications of such high performance across these five specific targets.
Lalam: Considering the potential impact, this work could help us understand which specific sequence features are most critical for a transcription factor binding site, which could inform how we design stronger regulatory elements. It’s about making the design process more informed rather than purely empirical.
Tom: So, to wrap up this summary of "From DNA Design to DNA Slimming: Auditable Agentic Discovery of a Deletion-Only Designer," we've seen that ERA discovered GRADASLIM, which is a deletion-only sequence designer that consistently beats random and greedy methods across five targets.
Jane: And the authors are Joel Shor, at the Allen Institute and Move37 Labs. The implication is that we now have a machine-verifiable method for creating compact regulatory DNA by removing unnecessary parts while preserving function.
Lu: I think the real weight of this paper lies in establishing deletion-only slimming as a distinct, machine-verifiable design problem, which provides a clear framework for future sequence optimization research.
Meng: From an engineering viewpoint, if the constraints on exact length and deletion only can be maintained with this level of performance across these targets, it opens up possibilities for streamlining vector payloads in molecular biology.
Lalam: I believe the future impact is in using this to build more efficient biological systems, where we can reduce the size of our genetic tools while ensuring they still perform their intended regulatory roles.
Tom: That’s what we’ve been talking about—a smarter way to design DNA by focusing on what to keep, not just what to swap in, thanks to this agentic discovery.
Conclusion: Tom: So, we've been diving deep into how ERA discovered GRADASLIM, this deletion-only sequence designer that’s outperforming other methods for five transcription factor targets on arXiv today.
Jane: It really is impressive to see how they managed to build a system that not only finds a sequence but actually proves it meets those strict, machine-verifiable deletion requirements.
Lu: What I find most intriguing is the framework of agentic discovery itself; it’s like we're teaching an AI how to be a meticulous editor for DNA, constrained by hard rules.
Meng: From my side, I'm still thinking about the computational overhead involved in maintaining that salience vector across such iterative refinement steps. How scalable is this really when you look at larger sequence sets?
Lalam: I see this as a major step forward because it shows we can move toward designing biological elements not just by guessing, but by having an AI systematically search for and validate the exact right structural modification.
Tom: Exactly, Lalam. And the authors, Joel Shor from Allen Institute and Move37 Labs, have really laid out how this works in a way that makes it transparent to us.
Jane: It gives us a clearer path for understanding *why* certain sequences are effective by showing exactly which features drive that predicted activity.
Lu: The implications for sequence analysis are huge; it suggests a new way to interpret the regulatory landscape without relying solely on traditional, brute-force searching techniques.
Tom: And I think the real excitement here is seeing how this capability can translate from a lab concept into actual tools that help us design more efficient genetic circuits.
Jane: It opens up a whole new avenue for improving the size and complexity of our biological tools by removing unnecessary parts with precision.
Meng: I'm looking forward to hearing more about the practical constraints they faced during those evaluations, specifically around those different target lengths they tested.
Lalam: I think this work could fundamentally improve how we approach sequence optimization, offering a more systematic and auditable method for creating functional DNA sequences.
Joel Shor
Allen Institute & Move37 Labs
cs.LG, cs.NE, q-bio.GN
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: 40th Conference on Neural Information Processing Systems (NeurIPS 2026). Workshop: Agentic AI for Biological Discovery
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 74/100
The gist: Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity.
Key concepts
- Sequence Slimming
- This is the task of selecting a shorter DNA subsequence from a larger one such that the new sequence maintains its original biological activity. The goal is to find an exact-length, order-preserving subsequence where only deletions are allowed, minimizing energy loss while maximizing retained effect.
- Agentic Discovery
- This involves using an AI agent (ERA) to search for a solution by modifying existing programs. ERA was tasked with turning a starting program into a specific sequence designer by searching over possible DNA subsequences and applying constraints to ensure only deletion-only changes are made.
- GRADASLIM Algorithm
- This is the discovered two-phase search strategy used by ERA. Phase 1 quickly finds candidates by deleting bases from 'low-salience' positions. Phase 2 refines these candidates using local modifications like swaps and shifts, guided by a 'salience vector' that tracks which sequence positions are most important for retaining activity.
- Salience Vector
- This is a dynamic system within GRADASLIM that acts as a memory of importance rather than just a mathematical gradient. It credits positions in sequences that lead to better-than-average candidates and spreads this credit to neighboring bases, helping the algorithm preserve important motifs during refinement.
Terminology
Summary
Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity. The gist: ERA discovered a deletion-only sequence designer that outperforms random and greedy methods for five transcription-factor binding targets.
Problem Definition
The task of sequence slimming
is defined as selecting an exact-length, order-preserving subsequence while retaining activity.
This involves finding a candidate sequence that equals an indexed subsequence from the source DNA, where the indices must be strictly increasing and within the original sequence bounds. The optimization problem seeks to minimize energy by maximizing the retained effect, quantified by the metric R: R = E(b) − E(Embed(xI, b)). (2)
. This formulation ensures that only deletions are permitted; substitutions, insertions, reordering, or inversion fail a machine-verifiable reconstruction check.
Agentic Discovery of the Designer
The research utilized an Agentic program search approach where the Empirical Research Assistant (ERA) searched over executable designer programs. ERA received a starting program, GrAdaBeam (a hybrid of gradient guidance and adaptive discrete search), and was tasked with modifying it to produce a deletion-only sequence slimmer. The process involved nested problem decomposition,
where the outer loop searches for ordinary code that can be inspected, while each proposed program searches over DNA subsequences. ERA constrained itself using hard checks
to prevent cheating the deletion-only requirement, prioritizing reaching the length constraint before optimizing activity.
The Discovered Algorithm (GRADASLIM)
The best program discovered by ERA is named GRADASLIM, which employs a two-phase search strategy. Phase 1 focuses on reaching the length constraint quickly by sampling deletions preferentially from low-salience positions.
Phase 2 involves exact-length refinement,
utilizing three types of local source geometry modifications: salience-biased swap,
point shift,
and contiguous block shift.
The algorithm updates a persistent salience vector, which is described as a reservoir of credit assignment rather than a model gradient. This salience is updated by crediting positions retained in better-than-average candidates and diffusing small amounts of credit to neighboring positions.
Evaluation Results
The discovered program was evaluated across five targets: E2F3, ELF4, MAX, MECOM, and RAD21, at target lengths of 400 bp and 100 bp. The results show that ERA has the highest mean in 9/10 settings
and exceeds random deletion in all ten.
Specifically at 400 bp, ERA outperforms greedy for every target, with paired bootstrap intervals above zero for every target-specific comparison. At 100 bp, ERA is numerically higher for ELF4, MAX, MECOM, and RAD21. The study concludes that ERA’s slimmed sequences improve the predicted lift over background relative to their full-length sources,
as the mean retained effect exceeds 1 in every setting.
Limitations and Conclusions
The evaluation design has limitations, including testing only two output lengths and using fixed, target-specific contexts. Furthermore, ERA program search did not include 100-bp targets, meaning the 100-bp evaluation tests substantially more aggressive slimming than was used during algorithm discovery.
The frequent R > 1 values observed may reflect exploitation of the BPNet oracle or artificial deletion junctions rather than preservation of biological function.
Nevertheless, this work establishes deletion-only slimming as a distinct, machine-verifiable design problem, showing a consistent ERA advantage at 400 bp and mixed performance at 100 bp.
Algorithm Summary (Algorithm 1)
The algorithm iteratively refines the sequence by cycling between Phase 1 (reaching the constraint quickly) and Phase 2 (exact-length refinement). In Phase 2, it samples from a set of candidate modifications—swap, point shift, or block shift—based on probability thresholds. The core mechanism involves updating the salience vector based on performance: positions occurring in below-average-energy candidates gain salience.
Neighbor diffusion biases nearby bases together to preserve motifs. Simulated annealing allows the single incumbent to cross local barriers while an archive preserves the best valid candidate found so far. The process continues until a solution of exact length is found that minimizes energy relative to the source sequence.
The gist
ERA discovered a deletion-only sequence designer that outperforms random and greedy methods for five transcription-factor binding targets. ERA has the highest mean retained effect in nine of ten settings and exceeds random deletion in all ten. At 400 bp, ERA outperforms greedy for every target and all 25 paired starts, with every target-specific interval above zero. At 100 bp, ERA is numerically higher for ELF4, MAX, MECOM, and RAD21.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems, particularly those involved in DNA/nucleic acid design:
-
Enhance sequence design capabilities by incorporating a
sequence slimming
module based on the ERA-discovered GRADASLIM algorithm. -
Develop a machine-verifiable evaluation framework specifically for deletion-only sequence slimming tasks, ensuring that AI outputs are rigorously checked against exact length and order constraints before being accepted as valid candidates.
-
Improve the efficiency and performance of model-based nucleic acid designers by shifting from purely substitution-based optimization to incorporating deletion actions, leveraging hybrid search architectures like GrAdaBeam (Gradient guidance + Adaptive discrete search).
-
Implement an agentic program search framework (similar to ERA) that can autonomously discover novel design strategies (like the GRADASLIM algorithm) by searching over executable designer programs rather than just optimizing fixed-length sequence modifications.
-
Create a salience map mechanism within the AI's objective function that learns which source positions are most critical for maintaining predicted activity, allowing the system to prioritize deletions of low-salience bases while preserving high-salience ones during sequence compaction.
These improvements will enable the resulting AI systems to:
-
Design more compact viral vector payloads by intelligently removing non-essential regulatory elements while preserving the binding function of transcription factors.
-
Generate novel, functional DNA sequences that are provably shorter than existing optimized designs, leading to reduced synthesis and assay costs in wet-lab settings (e.g., transcription factor binding studies).
-
Perform sequence optimization that is more robust and less prone to local optima by combining gradient-based refinement with targeted discrete search strategies for exact length constraints.
-
Autonomously discover high-performing, specialized design algorithms tailored to specific biological constraints, such as the deletion-only problem, without requiring manual algorithmic engineering.
Abstract
Compact regulatory DNA can free up space in vector payloads, reduce synthesis and assay burden, and expose which sequence features drive predicted activity. Yet most model-based nucleic-acid designers optimize fixed-length sequences through substitutions; they do not ask which bases of an existing functional element can be removed while retaining predicted activity. We define the task of sequence slimming as selecting an exact-length, order-preserving subsequence while retaining activity. Modeled on the design benchmark NucleoBench, we propose a quantitative evaluation for slimming that balances sequence reduction with maintaining function. Each slimmer must return both the subsequence and its source indices, which can be used to verify that the slimmer obeyed task requirements. To our knowledge, this is the first dedicated benchmark of this deletion-only problem. The coding agent Empirical Research Assistant (ERA) then searched over executable designer programs. ERA received the task prompt and a successful substitution-only designer GrAdaBeam as a starting program, and it modified the designer to produce GRADASLIM. We report held-out evaluations for five transcription-factor binding targets, comparing random, greedy, and ERA-guided slimming at 400 and 100 bp. ERA has the highest mean in 9/10 settings. Paired bootstrap intervals for ERA minus greedy are above zero in all five 400-bp settings, below zero in one 100-bp setting, and overlap zero in the remaining four.
Sources
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- AdaLead: A simple and robust adaptive greedy search algorithm for sequence design
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks