DarwinX: Evolving Agent Harnesses Through Natural Selection

arXiv:2608.07545 · cs.NE, cs.AI, cs.LG, cs.SE · Submitted 2026-07-31 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DarwinX: Evolving Agent Harnesses Through Natural Selection".

Jane: The paper was written by Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur et al. from Salesforce AI Research and Salesforce Agentforce.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re looking at a paper that’s been making waves in the agent research world, and it’s called “DarwinX: Evolving Agent Harnesses Through Natural Selection.” Jane, I have to say, just the title alone got me excited.

Jane: Same here, Tom. And for anyone just tuning in, let’s break down what that title actually means. An “agent harness” is basically everything around the AI model itself — the prompts, the tools, the memory, the control flow. Think of it as the cockpit the model sits in. And “evolving through natural selection” means they’re not hand-designing that cockpit; they’re letting versions of it compete and only the fittest survive.

Tom: Exactly. And what’s wild is that the model inside the cockpit stays completely frozen. No weight updates, no retraining. All the improvement comes from changing the harness around it. That’s a huge philosophical shift, right? We usually think the model is the only thing that matters.

Jane: Right, and that’s why this paper is so important. It’s from Salesforce AI Research, and the team includes folks like Yifan Zhang and Yutong Dai as first authors, with senior guidance from Silvio Savarese and others. They’re basically saying: stop treating the model as the only lever you can pull.

Tom: And the results back that up. They took a frozen model and improved its performance on Terminal-Bench two point one from seventy-five point five percent to eighty-three point two percent — that’s a seven point seven point jump just by evolving the harness. And on WebArena-Infinity, they went from forty-three point five percent to ninety-three percent audit-clean. I mean, that’s almost double.

Jane: It really makes you wonder how much untapped capability is sitting in models we already have, just waiting for a better harness to unlock it. And that’s the hook for today — we’re going to dig into how this “natural selection” actually works, and what it means for the future of building agents.

Tom: Stick around, because this one’s a game changer.

Summary: Jane: So Tom, we’ve set the stage with the title. Now let’s talk about what the paper actually does, because the summary is pretty dense. At its core, DarwinX treats self-improvement as a selection problem over a population of harnesses, not a training problem.

Tom: And that distinction matters. They’re not saying “let’s fine-tune the model.” They’re saying “let’s breed better harnesses.” Each harness is like an organism. It gets mutated — small edits to prompts, tools, control flow — and then it’s tested. If it solves something new without breaking what it already solved, it survives. If not, it gets reverted.

Jane: But here’s the clever part. They don’t just keep the winners. They keep everything in an archive, even the losers. Because a variant that’s worse overall might still be the only one that cracked a particular task. And later, they can merge complementary specialists — like combining the best traits of two different lineages — to get a harness that’s better than either one alone.

Tom: That’s the “population” part of the title. And the “natural selection” part is the preserve-and-extend contract. Every child harness has to earn its place. It has to show a net gain on some tasks while only regressing a tiny bit on others. No gold solutions, no hand-picked winners — just measured fitness under the benchmark’s own verifier.

Jane: And the results across four benchmarks are honestly stunning. We mentioned Terminal-Bench and WebArena. They also tested on TerminalWorld, a held-out split, where they hit sixty-eight point three percent — beating every off-the-shelf agent they compared against. And then they took the Terminal-Bench harness and ran it unchanged on SWE-bench Verified, hitting eighty-four point two percent without any in-domain feedback.

Tom: So the harness learned general competence, not benchmark-specific tricks. That’s the headline. It transfers across tasks, across verifiers, and even across base models. Lu, you’ve been quiet — what do you make of this from a research perspective?

Lu: I think the most exciting implication is that we’ve been leaving capability on the table. The model is frozen, but the harness is a vast search space we’ve barely explored. This paper shows that space is rich enough to double performance in some cases. That reframes where the frontier of agent research actually is.

Tom: And that’s exactly where we’re headed next — the mechanism behind this evolution. But first, let’s just sit with that summary. A frozen model, a better harness, and a fifty-point jump on a real-world benchmark. That’s not incremental. That’s a new paradigm.

Improvements: Jane: Welcome back. We’ve covered the big picture, and now I want to get into the specific improvements this paper proposes over prior work. Because DarwinX didn’t invent self-improving agents — there were already systems like SICA and the Darwin Gödel Machine. But this paper fixes two big failure modes.

Tom: Right, and those failure modes are path dependence and cross-task interference. Path dependence means if you only follow one lineage of edits, you get stuck on early decisions. Cross-task interference means an edit that helps one type of task silently breaks another. Both of those kill single-lineage self-editors.

Jane: So DarwinX’s improvement is the archive and the recombination. Instead of one line of descent, you have a tree of harnesses. You keep alternative lineages alive. And when you find two specialists that solve different tasks, you merge them. The merged child only survives if it inherits both parents’ wins.

Tom: And that’s a real improvement over something like the Darwin Gödel Machine, which mutates one parent at a time and scores the child against that parent. It never brings complementary lineages back together. DarwinX does. It’s like evolution with sexual reproduction instead of just cloning.

Lu: And there’s another improvement I want to highlight. The proposal signal is modular. They have three types of evidence that can drive edits: failure-derived, teacher-derived, and self-derived. So if the agent keeps failing a task, the system analyzes those failures. If there’s a stronger solver’s successful trajectory, it distills that. And if the agent has both passes and fails on the same task, it contrasts them.

Meng: From an engineering standpoint, that modularity is huge. It means you can plug in new evidence sources without rewriting the whole loop. And the two-speed selection — promote on bounded risk, then confirm on stricter avg@k — that’s a practical answer to noisy benchmarks. You don’t let one lucky rollout redirect the search.

Jane: Meng, that’s a great point. The paper is very explicit that on noisy agent benchmarks, a single lucky rollout can masquerade as a capability gain. So they separate exploration from confirmation. A variant can enter the tree on promising but noisy evidence, but it only earns the right to steer future search after clearing a stricter re-test.

Tom: And that’s the improvement that makes the whole thing work. It’s not just “try stuff and keep what scores higher.” It’s “try stuff, keep what survives re-verification, and preserve what you already solved.” That’s how you compound gains instead of trading one task for another.

Jane: Exactly. And speaking of compounding, the next segment is going to look at the actual first page of the paper, where they lay out this vision in their own words. That’s where the ambition really shows.

First Page: Tom: So we’re now looking at the opening of “DarwinX: Evolving Agent Harnesses Through Natural Selection,” and honestly, the first page reads like a manifesto. They open by saying an agent’s capability depends not just on model weights but on its harness — the prompts, tools, skills, and control flow.

Jane: And then they drop this line that I love: “A frozen model need not be a fixed agent.” That’s the thesis. You don’t have to retrain to get better. You can evolve the harness around the model, and that turns evaluation compute into durable capability.

Tom: Durable capability. That’s the key phrase. Because when you evolve a harness, the gains persist. They’re not baked into weights that might get overwritten. They’re in the prompts and skills and control flow, which you can keep, copy, and transfer to other models.

Lu: And that’s why the first page is so important. They’re positioning this as natural selection, literally. No gold labels, no hand-picked winners. Only survival of the fitter variant under measured fitness. That’s a strong claim, and they back it up with the preserve-and-extend contract.

Meng: I also appreciate that the first page is honest about the failure modes they’re addressing. They cite Robeyns et al. reporting that single-lineage self-editors plateau. And they cite the Darwin Agent Team noting that isolating variants contains interference but leaves specialists in separate lineages. DarwinX is a direct response to both.

Jane: And the first page also previews the evaluation ladder, which I think is really smart. Four benchmarks ordered by increasing separation between the evolution signal and the test. Terminal-Bench is in-domain. TerminalWorld is held-out tasks. WebArena-Infinity is synthetic-to-real. And SWE-bench is cross-benchmark transfer.

Tom: That ladder is what makes the results convincing. Because if you only improved on the benchmark you optimized for, you’d have a glorified overfitter. But they improve across all four, with the model frozen. That’s general agent competence, not benchmark-specific patches.

Lu: And that’s the vision that gets me excited. We’re not just making better agents for benchmarks. We’re building a process that makes better agents, period. The harness is the substrate, and selection is the algorithm. That could apply to any domain where you can measure fitness.

Meng: And from a deployment standpoint, the auditability is huge. Every harness edit is human-readable. You can see exactly what changed and why. That’s a property weight-space self-improvement doesn’t have.

Jane: So the first page sets up a big promise. And in the conclusion, we’re going to zoom out and ask: what does this mean for the world? What happens when harness evolution becomes standard practice?

Conclusion: Tom: And we’re back for the final segment on “DarwinX: Evolving Agent Harnesses Through Natural Selection.” Jane, we’ve covered the title, the summary, the improvements, and the first page. Let’s pull it all together.

Jane: So the core message is that a frozen model is not a fixed agent. By evolving the harness — the prompts, tools, skills, and control flow — you can unlock massive capability gains without touching the weights. And the results speak for themselves: Terminal-Bench up seven point seven points, WebArena-Infinity up forty-nine point five points audit-clean, TerminalWorld beating every off-the-shelf agent, and clean transfer to SWE-bench.

Tom: And the mechanism is natural selection over a population of harnesses. Preserve-and-extend contract, archive of alternative lineages, recombination of complementary specialists. No gold solutions, no hand-picked winners. Just measured fitness under the benchmark’s own verifier.

Lu: What excites me is the broader implication. This reframes where agent capability comes from. We’ve been obsessed with bigger models. But this paper shows the harness is a vast, under-explored search space. And because harnesses are human-readable and transferable, they’re an asset that outlives any single model generation.

Meng: From a practical standpoint, the two-speed selection is what makes this deployable. You don’t let one lucky rollout redirect the search. You confirm on stricter measurement before a variant steers anything. That’s how you get reliable gains in noisy real-world environments.

Lalam: And if I may add, the cultural impact is significant. This is a shift from training to selection. It means smaller teams with limited compute can still build powerful agents by evolving harnesses on frozen models. It democratizes agent improvement. And the auditability means we can trust what evolves, because every change is a readable diff, not a black-box weight update.

Jane: That’s a beautiful way to put it, Lalam. And it’s the note we want to end on. DarwinX isn’t just a paper about benchmarks. It’s a paper about a new way to think about intelligence — not as something you train, but as something you cultivate.

Tom: And with that, we’re saying goodbye to “DarwinX: Evolving Agent Harnesses Through Natural Selection.” Thanks for listening, and we’ll see you next time with another paper that’s pushing the frontier.

Jane: Take care, everyone.

Yifan Zhang, Yutong Dai, Juntao Tan, Luyu Yang, Rishi Mullur, Thai Hoang, Zhiyuan Hu, James Zhu, Phil Mui, Silvio Savarese, Ran Xu, Zeyuan Chen

Salesforce AI Research · Salesforce Agentforce

cs.NE, cs.AI, cs.LG, cs.SE

Submitted: 2026-07-31

Updated: 2026-08-11

Code: https://github.com/browser-use/browsercode

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 63/100

The gist: The paper introduces DarwinX, a system that treats self-evolution of LLM agents as "selection over a population of harnesses with the model frozen." The central claim is that "an LLM agent's

Terminology

Summary

The paper introduces DarwinX, a system that treats self-evolution of LLM agents as selection over a population of harnesses with the model frozen. The central claim is that an LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow, and that a frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.

Core mechanism. DarwinX operates under a "preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. The system has three parts: (1) a branch evolution loop with repeated trace-guided edits and survival selection: candidates with bounded regression risk may become ancestors, while stricter avg@k confirmation controls promotion; (2) a population-level inheritance loop where surviving and archived branches form a population; complementary specialists can be inherited or merged so that improvements discovered in different lineages are not trapped apart; and (3) a modular learning-signal loop where failure-derived diagnosis, teacher-derived demonstrations, or self-derived rollout contrast are converted into harness edits rather than model-weight updates. Selection is driven purely by measured fitness: a variant's avg@k solve rate under the benchmark's own verifier, with no gold solutions and no hand-picked winners."

Fitness and promotion. Each variant is scored by per-task solve rate p̂t(v) (avg@k). For a child c and parent p, the per-task change Δt = p̂t(c) − p̂t(p) is summarized as net gain g(c) = Σt Δt and bounded regression R(c) = Σt (−Δt)+. "The fitness enabler admits a child that extends without breaking preservation, i.e. g(c) > 0 and R(c) ≤ δ. A reasoned verifier agent f adjudicates in two stages (promote, then probe), and a promoted child is re-tested at higher fidelity with a preservation probe before it may steer search. Each node carries a lineage gain G(c) = G(p) + g(c). Parent selection ranks nodes by cumulative lineage gain, sampling p∗ ∼ (1 − β) δarg maxv∈S G(v) + β Broaden(P)" — exploiting the highest-gain confirmed node with probability 1−β, otherwise broadening across the population.

Population and recombination. Every scored variant is retained as a first-class archive node. Variants are classified by how their solved set S(v) compares to the parent's: "improvers (S(c) ⊋ S(p)) and neutral children (S(c) = S(p)) preserve every inherited solve and stay eligible for inheritance. The rest give one up (S(c) ̸⊇ S(p)) and are not eligible, feeding back only their distilled lessons. When variants v1,..., vn solve complementary tasks, DarwinX materializes an inherited child by merging additive edits: from a common ancestor H0, the merged harness is H = H0 ⊕ Δ with Δ = Δcode ⊕ Δskill ⊕ Δprompt ⊕ Δtool, and the child is kept iff it covers the union of its parents' wins, S(child) ⊇ ∪i S(vi)."

Learning signals. Three signal types are used: "Failure-derived signals (∇) summarize failed trajectories τ and localize missing capabilities. Teacher-derived signals (π∗) distill a reference solver's successful trajectory τ∗ into a reusable approach. Self-derived signals (A) contrast the agent's own passing and failing rollouts τi ki=1 to identify what makes success reliable. Failure-derived signals are the default for ordinary mutations; teacher-derived signals enrich the analyzer on walls (tasks with no successful agent rollout); self-derived signals enrich the analyzer on variance-band tasks" where both passing and failing rollouts exist.

Evaluation design. The paper evaluates across four benchmarks "ordered by increasing separation between the evolution signal and the test: in-domain test-time evolution (Terminal-Bench 2.1), held-out task generalization (TerminalWorld), synthetic-to-real generalization (WebArena-Infinity), and cross-benchmark transfer (Terminal-Bench 2.1 → SWE-bench Verified)." A fifth question is answered by an ablation with case studies.

Results on Terminal-Bench 2.1 (RQ1). On a frozen GPT-5.5 base, DarwinX lifts base Monet from 75.5% to 83.2% avg@5 (+7.7 points) under the strict leaderboard protocol. On a frozen GPT-5.6 Sol at medium effort, DarwinX scores 84.7%, at the frontier of the verified Terminal-Bench 2.1 leaderboard: it matches or exceeds the current verified leader (Claude Code + Fable 5, 83.8% at xhigh) while running at a lower effort setting, and adds +2.9 points over OpenAI's own native single-agent GPT-5.6 Sol at the same medium effort (81.8%). Gains concentrate in ML & scientific-computing tasks (60.1 → 74.9%, +14.8 points) and data/database tasks (83.9 → 97.8%, +13.8). The paper argues the gain is not compute: the extra compute DarwinX spends is productive and targeted: it concentrates on exactly the tasks it newly solves — on the six tasks that flip from failing to solved, the evolved harness roughly doubles turns (22 vs. 11) and quadruples tokens (380K vs. 89K), while on the 69 tasks both agents already solve, compute barely moves (13 vs. 12 turns). The reward-hacking audit found no harness-level cheating, with only two of 370 rewarded trajectories flagged, one a false positive and one a single confirmed shortcut (an mteb-leaderboard trial where the agent read an answer-bearing string from the task's own published README), which does not alter the aggregate conclusion.

Results on TerminalWorld (RQ2). On a frozen Opus 4.8 base, the harness evolves on 94 training tasks and is frozen before single-attempt pass@1 evaluation on 41 disjoint held-out tasks. Monet (DarwinX) on Opus 4.8 resolves 28/41 (68.3%), the best result on the split and above every off-the-shelf agent we evaluate, including Claude Code (65.9%), base Monet (61.0%), and Terminus-2 (61.0%). The matched GPT-5.5 pair moves 20 → 23. The paper notes the in-loop proxy overfits: the training-subset score saturates from 0.505 to 1.000, yet held-out pass@1 is 68.3%: a 31.7-point gap. Crucially, "the variant that best fits the proxy is not the best generalizer. Four high-scoring specialist variants solve 24, 25, 26, and 27 of the 41 held-out tasks on overlapping but distinct subsets, and the merged harness reaches 28, above every individual specialist. The paper cautions that with only 41 tasks, one solve moves pass@1 by 2.4 points, and we treat the one-task margin over the strongest off-the-shelf agent (Claude Code, Opus 4.8) as suggestive rather than statistically decisive (paired exact McNemar p = 1.0)."

Results on WebArena-Infinity (RQ3). On a frozen GPT-5.5 base, evolution operates on 300 synthetic intents scored by an LLM judge; reporting uses deterministic pass@1 on the 1,260 unseen real tasks. Monet (DarwinX) achieves the best audited result at 93.0% audit-clean pass@1, outperforming GPT-5.5 + Browser Use (86.1%) by 6.9 points and the top public agent, Gemini 3 Flash + Browser Use (69.3%), by 23.7 points. Relative to base Monet on the same frozen GPT-5.5, the evolved harness improves from 43.5% to 93.0% audit-clean (+49.5 points). The gain is broad: every application improves, with the largest gains on state-change-heavy applications (Elation prescriptions +75.0, Gmail +73.3, Gmail accounts/contacts +70.0). The validity audit shows capability and compliance improve together: "Base Monet has 155 evaluation-plane violations, 97 privileged-host violations, 26 exploit or privilege-escalation violations, and 15 raw-state mutations. The first three classes disappear after evolution; the remaining 17 violations of the evolved harness are raw-state mutations. Evolution reduces invalid trajectories from 293 to 17."

Results on SWE-bench Verified transfer (RQ4). On a frozen Opus 4.8 base, the best Terminal-Bench 2.1 harness is run unchanged on all 500 SWE-bench Verified issues and graded by the official test harness (pass@1). The TB2.1-specialized harness reaches 421/500 (84.2%) official pass@1, +3.4 points over the 80.8% fix-skill reference, without receiving any SWE-V feedback. The paper scopes this: SWE-V is a transfer target only... The in-loop signal available for this benchmark scored trajectory completion rather than official test resolution, so it is not a sound basis for selection here.

Ablation (RQ5). On Terminal-Bench 2.1, the evolved lineage adds seven harness skills, and every one belongs to a single family, verification / artifact-contract: verifier-contract, contract-candidate, graded-artifact-final-check, artifact-verification-loop, real-tool-artifact, tool-grounded-artifact, and security-contract-repair. None adds domain knowledge; each makes the agent establish and check an explicit acceptance contract, or ground its output in real tool execution, before finalizing. The paper notes this is an exploratory attribution, not a per-skill causal ablation: the skills were co-selected, not independently randomized.

Limitations. The paper acknowledges: The experiments evaluate the complete DarwinX system: the archive, parent selector, recombination operator, and inference effort are not independently randomized; TerminalWorld's matched Opus comparison is suggestive rather than decisive; SWE-V official scores span just 80.8–84.2%; and the WAI audit is far stronger than a keyword heuristic, but it is not a formal sandbox.

Overall conclusion. "Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:


  • What I add: A persistent archive of harness variants (prompts, tool definitions, control flow, skill documents) stored as a tree, where each node records its edit delta, per-task scores, trial evidence, and distilled lessons. Selection uses a preserve-and-extend contract: a child variant is admitted only if it improves at least one task (g(c) > 0) and regresses no more than a small tolerance (R(c) ≤ δ). Promotion is two-stage: a permissive fitness enabler admits bounded-risk gains, then a stricter avg@k re-test and preservation probe confirm before the variant may steer future search.

  • What the improved system can do: It no longer hill-climbs from a single lineage. It retains losing variants as genetic material, recombines complementary specialists across branches, and compounds improvements without trading one capability for another. It can recover from early bad edits and avoid local plateaus.

  • What I add: Three interchangeable evidence sources that all feed the same edit-proposal loop: (a) failure-derived signals (summarize failed trajectories, localize missing capabilities), (b) teacher-derived signals (distill a reference solver's successful trajectory into a reusable approach), (c) self-derived signals (contrast the agent's own passing and failing rollouts on the same task). The system dynamically partitions tasks into reliable solves, variance-band tasks, and walls to choose the most informative signal per task.

  • What the improved system can do: It can improve even on tasks where it has zero successful rollouts (using teacher signals) and on tasks where success is flaky (using self-contrast). It never proposes edits blind; it always has the most informative evidence available.

  • What I add: A set of harness skills that force the agent to (a) derive an explicit acceptance contract from the task description, (b) check the candidate solution against that contract before finalizing, (c) ground outputs in real tool execution rather than asserted results, and (d) iterate a fix-and-recheck loop until the contract is satisfied. These are implemented as prompt-level skills and control-flow modifications, not model-weight changes.

  • What the improved system can do: It reduces wrong-output and premature-finalization errors. On Terminal-Bench 2.1, this alone accounts for most of the +7.7-point gain (75.5% → 83.2%), concentrated on ML/scientific-computing and data/database tasks where procedural verification was the bottleneck.

  • What I add: Before any promoted child becomes an ancestor, the system re-samples the lineage's known solved-task set at higher fidelity (avg@5 instead of avg@3) and requires that no previously-solved task regresses beyond a noise band. This is enforced at every generation, not just at final evaluation.

  • What the improved system can do: It prevents the classic self-improvement failure where an edit fixes one task family while silently breaking another. The system can safely evolve on mixed task distributions without stagnation, because the selection criterion reflects the full benchmark rather than a narrow local objective.

  • What I add: A merge operator that takes two or more complementary specialists (each solving a different subset of tasks), combines their additive edits from a common ancestor, and accepts the merged child only if it covers the union of both parents' solved-task sets (S(child) ⊇ S(v1) ∪ S(v2)). The source pool includes archived losers that each contribute a unique solved task.

  • What the improved system can do: It can combine capabilities that no single lineage ever discovered together. In TerminalWorld, four specialists solved 24–27 held-out tasks each on overlapping subsets; the merged harness solved 28, beating every individual specialist and every off-the-shelf agent.

  • What I add: A failure-mode classifier labels each trial (e.g., timeout-setup, wrong-output, tool-error) and aggregates dominant themes across the benchmark. These themes are stored in a shared memory K g that both the proposer and verifier read. The dominant theme is injected into the proposal prompt so the search invents global capabilities (e.g., setup cost dominates timeouts → build an efficient-setup capability) rather than per-task patches.

  • What the improved system can do: It addresses systemic bottlenecks across the whole task distribution, not just individual failures. This is what allowed the WAI harness to improve from 43.5% to 93.0% audit-clean pass@1 across 10 unrelated applications.

  • What I add: A static analyzer (Stage 1) that de-obfuscates JavaScript, taint-tracks scored collections, and flags evaluation-plane access, privileged-host knowledge, raw-state mutation, database manipulation, and exploits. Stage 2 sends flagged trajectories to an independent LLM judge under the same rubric; unresolved cases go to human review. Invalid trajectories count as failures even if the verifier passes them.

  • What the improved system can do: It reports audit-clean scores that are robust to reward hacking. In WAI, this reduced invalid trajectories from 293 to 17 (all residual ones being raw-state mutations in a single application), and eliminated evaluation-plane access, privileged-knowledge, and exploit mechanisms entirely. Capability and compliance improve together, not at each other's expense.

  • What I add: The harness learns to spend additional test-time compute (turns and tokens) only on tasks where reasoning was previously insufficient, leaving already-solved tasks at their baseline compute. This is learned implicitly through the preserve-and-extend contract: edits that add compute to solved tasks regress nothing but also gain nothing, so they are not promoted.

  • What the improved system can do: It achieves higher scores without a uniformly larger budget. On Terminal-Bench 2.1, the evolved harness doubled turns (22 vs. 11) and quadrupled tokens (380K vs. 89K) on the six newly-solved tasks, while compute on the 69 already-solved tasks barely moved (13 vs. 12 turns). The gain comes from knowing when to keep working, not from spending more everywhere.

  • Self-improve without retraining: The base model stays frozen; all gains come from the harness. This means the system can be improved using only evaluation compute, with no GPU training budget.

  • Generalize beyond the evolution signal: It transfers across task distributions (terminal → browser), reward sources (LLM judge → deterministic verifier), and even across benchmarks (Terminal-Bench 2.1 → SWE-bench Verified, 84.2% without any in-domain feedback).

  • Improve capability and compliance together: The preservation contract and validity audit ensure that gains are durable and legitimate, not fragile verifier-gaming.

  • Recover from bad early edits: The archive retains alternative lineages, so a poor initial choice does not doom the search.

  • Combine complementary strengths: Recombination lets the system merge specialists that each solve different tasks, achieving coverage no single lineage could reach.

  • Scale to heterogeneous task distributions: The cross-task theme aggregator and preservation probe prevent the stagnation that plagues single-lineage self-editors on mixed benchmarks.

Abstract

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fitness comes from each benchmark's own verifier: no gold solutions, no hand-picked winners. Across four benchmarks that progressively separate the evolution signal from the test, one loop adds about 17 points on average: Terminal-Bench 2.1 rises +7.7 to 83.2% on a matched base and to the verified frontier at 84.7% on a stronger one; TerminalWorld's held-out split reaches 68.3%, ahead of every off-the-shelf agent; WebArena-Infinity real-task pass@1 rises from 43.5% to 93.0% audit-clean; and a Terminal-Bench 2.1 harness transfers unchanged to SWE-bench Verified. What evolves is general agent competence, not benchmark-specific patches, so it survives changes of task, verifier, and base model. A frozen model need not be a fixed agent: harness selection turns evaluation compute into durable capability.

Sources

Related papers