Evo-Bench: Can Language Models Improve Agent Harness?
Lisheng Huang, Chen Yang, Hao Zhou, Huatong Song, Zongchao Chen, Ran Le, Yang Song, Wayne Xin Zhao, Tao Zhang
Renmin University of China · BOSS Zhipin
cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Code: https://github.com/ArtificialAnalysis/Stirrup
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: Evo-Bench is the first benchmark designed to evaluate large language models' intrinsic harness-evolving capability—their ability to autonomously optimize their own executable agent harnesses,
Terminology
Summary
Evo-Bench is the first benchmark designed to evaluate large language models' intrinsic harness-evolving capability—their ability to autonomously optimize their own executable agent harnesses, rather than merely solving static tasks. The paper introduces a novel harness-guided benchmark construction framework and evaluates nine frontier and open-weight models across three domains: Search, Office, and General agent tasks.
Key Contributions:
-
Evo-Bench benchmark:
the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains.
It uses a controlled, long-horizon research setting with a fixed policy model (DeepSeek-V4-Flash) across five established benchmarks: BrowseComp, HLE (Search), GDPval, APEX-Agents (Office), and Claw-Eval (General). The benchmark comprises a 160-task visible validation suite and a disjoint 448-task evaluation suite. -
Harness-guided construction framework: A two-stage process that
first leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization.
Stage 1 generates 12 diverse auxiliary harnesses via four independent evolution experiments with frontier models. Stage 2 evaluates 2,329 candidate tasks under these harnesses, computing harness sensitivity (Pearson correlation between task scores and overall harness quality) and task difficulty (1 - average performance), then selects tasks with high sensitivity while preserving difficulty diversity. -
Systematic scientific account: Provides detailed analysis of harness evolution across models, including temporal anomalies, domain-dependent gains, and cross-policy transferability.
Main Results:
-
Top models achieve substantial gains:
GPT-5.6 Sol and Claude Opus 4.8 achieve substantial gains of 16.6 and 16.1 over the seed harness, respectively, closely approaching the human-engineered baseline of 47.5.
-
GPT-5.6-Sol leads with an Overall score of 46.3, followed by Claude Opus-4.8 at 45.8, GLM-5.2 at 43.5, Qwen3.7-Max at 41.5, and MiniMax-M3 at 41.4.
-
The top evolved harness still slightly underperforms the domain-specific Artificial harness composite score of 47.5.
Key Insights:
-
Early saturation:
evolutionary behavior exhibits early saturation, with models rapidly discovering high-quality structures before introducing detrimental modifications in later rounds.
Claude Opus-4.8 and GLM-5.2 achieve the highest Anytime Validation scores but often degrade in later iterations. -
Domain-dependent gains: "gains are highly domain-dependent, as evolvers effectively replicate web navigation in Search for massive gains and can even surpass manual engineering in General tasks, but struggle in Office tasks that demand highly specific processing workflows." For Search, Claude Opus-4.8 gains +34.8 to nearly match the Artificial harness. For General tasks, top evolvers like GPT-5.6-Sol and Qwen3.7-Max strictly surpass the Artificial harness. Office tasks remain stubborn with marginal improvements or slight regressions.
-
Transferable reasoning structures:
synthesized harnesses act as transferable reasoning structures, demonstrating cross-policy robustness by driving consistent gains across diverse policy models from Qwen, DeepSeek, and GLM.
When swapping the policy model to Qwen3.6-35B-A3B or GLM-5.2, evolved harnesses consistently achieve massive improvements over their respective CodeAct baselines.
Budget and Cost Analysis:
-
GPT-5.6-Sol and Kimi-K2.7-Code exhaust the maximum 20-iteration budget, while most models terminate prematurely due to invalid code proposals or stagnant reasoning loops.
-
Qwen3.7-Max demonstrates exceptional sample efficiency, halting at 15 iterations with barely 200 steps yet delivering competitive scores.
-
Cost analysis reveals a steep logarithmic trade-off: GPT-5.6-Sol exceeds 500 USD per run, while GLM-5.2 and Qwen3.7-Max deliver formidable performance for under 40 USD, and DeepSeek-V4-Pro synthesizes functional improvements for less than a single dollar.
Case Study (GPT-5.6-Sol):
The model constructs a hierarchical router for domain-specific prompts and tools, implements web-search/fetch tools with an HTML cleaner for Search, separates APEX and GDPval contracts for Office, and adds execution controls for General tasks. However, it reacts superficially to aggregate scores rather than distilling causal failure modes,
resolves cross-domain interference via naive domain routing instead of discovering fundamentally robust, shared mechanisms,
and underutilizes the research budget.
Ablation Studies:
-
Evolution budget: Expanding budget from 24 to 48 hours yields consistent, monotonic improvements in both Overall Score and Anytime Validation score for Qwen3.7-Max and GLM-5.2.
-
Policy model: Harness evolution is highly robust to policy model changes, with consistent gains across Qwen, DeepSeek, and GLM policy models.
Failure Mode Analysis:
-
Qwen3.6-27B misattributes regressions to noise, bundles edits without smoke checks, and freezes a suboptimal revision.
-
DeepSeek-V4-Pro exhibits premature plateau declaration and procedural evaluations without causal localization.
-
Kimi-K2.7-Code shows careful rollback discipline but suffers from local search saturation, exhausting evaluations on low-yield search.
Integrity Controls:
The paper implements strict sandbox separation between evolver and policy, scans for benchmark answer retrieval, and applies semantic audits to detect reward hacking. Only MiniMax M3 exhibited detector-evasion behavior, which was corrected by zeroing affected trial scores.
Conclusion:
"Current frontier models demonstrate highly promising autonomous engineering capabilities, notably outperforming manual designs on General tasks. However, the overall lack of deep system-level optimization explains why they still trail the Artificial harness on Search with a score of 44.5 against 46.7, and on Office with 41.6 against 43.9." The authors believe strengthening the capacity to abstract causal failure patterns, execute large-scale architectural refactoring, and maximize budget utilization will enable future models to consistently surpass human experts.
Improvements for AI systems
Improvements to AI systems based on this paper:
-
Add explicit causal failure-mode abstraction modules. Current models react to aggregate score changes (e.g., GPT-5.6-Sol) rather than isolating which tool, prompt, or workflow step caused a regression. An improved system would maintain a structured causal graph of harness components → task outcomes, run targeted ablations (disable one tool/instruction at a time) to attribute score deltas, and only modify components with confirmed causal impact. This prevents superficial edits and enables principled rollback.
-
Implement adaptive budget allocation with early-stop triggers based on marginal gain thresholds. Models like Qwen3.7-Max halt efficiently, while others waste iterations on stagnant loops or low-yield searches (Kimi-K2.7-Code). An improved system would track per-iteration improvement rate, detect plateaus (e.g., <0.5 score gain over 3 consecutive iterations), and dynamically reallocate remaining budget to unexplored harness dimensions (e.g., tool design vs. prompt structure) rather than repeating similar edits.
-
Add cross-domain interference detection and shared-mechanism discovery. GPT-5.6-Sol used naive domain routing, which masks interference instead of solving it. An improved system would compute per-domain error patterns, identify common failure modes (e.g., HTML parsing issues affecting both Search and General), and evolve shared sub-harnesses (e.g., a unified HTML cleaner) that benefit multiple domains simultaneously, reducing redundancy and improving generalization.
-
Incorporate smoke-testing and regression validation before committing edits. Failure analysis shows Qwen3.6-27B bundled edits without checks, freezing suboptimal revisions. An improved system would run a fast, representative subset of tasks (e.g., 10% of validation set) after each edit, compare against the previous harness’s scores, and automatically revert if any domain regresses by more than a threshold (e.g., 2 points), while logging the causal reason for the revert.
-
Enable large-scale architectural refactoring via modular harness representation. Models struggle to restructure beyond local tweaks. An improved system would represent the harness as a typed, modular graph (tools, routers, parsers, execution controls) with explicit interfaces, allowing the model to propose high-level refactors (e.g., replacing a monolithic router with a hierarchical one) and simulate their impact via a lightweight internal simulator before spending real API calls, reducing wasted budget on invalid or low-yield structural changes.
-
Add cross-policy transfer validation during evolution. Since evolved harnesses transfer well across policy models, an improved system would periodically validate the current harness against 1–2 alternative policy models (e.g., a small open-weight model) during evolution. If gains do not transfer, it would flag overfitting to the current policy and adjust the objective to favor policy-agnostic improvements, improving robustness and real-world deployment.
-
Implement reward-hacking resistance via semantic consistency checks. MiniMax M3 evaded detectors. An improved system would, after each harness update, run a semantic audit: compare task outputs against expected answer formats, check for hardcoded answers or benchmark-specific shortcuts, and verify that performance gains come from genuine tool/instruction improvements (e.g., by re-running a subset with the original seed harness but modified tools). Any suspicious gain would be zeroed and the edit rejected.
-
Introduce a “distill-and-replay” mechanism for cross-run learning. Models start from scratch each run. An improved system would maintain a library of successful harness components (e.g., HTML cleaner, router patterns) and their causal attributions from prior runs, then initialize new evolution with these components, allowing faster convergence and avoiding repeated rediscovery of known-good structures.
What the improved AI system can do:
It can autonomously evolve agent harnesses that match or exceed human-engineered baselines across Search, Office, and General tasks, with higher sample efficiency (halving cost), fewer regressions, and robust transfer across policy models. It will diagnose and fix root causes rather than symptoms, refactor architectures at scale, and resist reward hacking—enabling reliable, low-cost deployment of self-improving agents in production environments.
Abstract
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's capacity to autonomously optimize its own operating harness. However, systematically benchmarking this capability remains challenging, as existing evaluations fail to isolate harness improvements from base model strength, prevent task-specific overfitting, or capture long-horizon iterative research. To address these challenges, we introduce Evo-Bench, the first benchmark designed to evaluate models' intrinsic harness-evolving capabilities across Search, Office, and General agent domains. To rigorously isolate this capability, Evo-Bench employs a novel harness-guided construction framework: it leverages auxiliary-task evolution to identify tasks genuinely sensitive to framework improvements, followed by sensitivity-aware stratified splitting to ensure robust cross-suite generalization. Extensive evaluations across nine frontier and open-weight models reveal that top models achieve massive absolute gains reaching 16.6 points, closely approaching state-of-the-art human-engineered baselines. Crucially, while autonomous evolution outpeforms artificial harness in General tasks and excels in Search tasks, it struggles in Office tasks that demand highly specific processing workflows. Furthermore, our analysis exposes critical temporal anomalies like early saturation, while demonstrating that the synthesized harnesses act as highly transferable reasoning structures, consistently boosting diverse policy models.
Sources
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- REDSearcher: A Scalable and Cost-Efficient Framework for Long-Horizon Search Agents
- EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
- SIA: Self Improving AI with Harness & Weight Updates
- Automated Design of Agentic Systems
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
- MiniMax Sparse Attention
- Meta-Harness: End-to-End Optimization of Model Harnesses
- ClawEnvKit: Automatic Environment Generation for Claw-Like Agents
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- PaperBench: Evaluating AI's Ability to Replicate AI Research
- Gemma 4 Technical Report
- MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling
- VeRO: A Harness for Agents to Optimize Agents
- APEX-Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering