HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
cs.SE, cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Comments: Project page: https://self-developing-agents.github.io/
Code: https://github.com/browser-use/browser-use
Project page: https://self-developing-agents.github.io
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness.
Terminology
Abstract
As agents move from research prototypes to deployed tools, their capability increasingly depends on model-external execution infrastructure, commonly termed the agent harness. Changing this harness while holding model weights fixed can substantially alter task performance. Current agent evaluations typically report downstream performance under a chosen harness, leaving a model's ability to develop the harness itself comparatively underexplored. We introduce HarnessDev, a benchmark that shifts the unit of evaluation from task outputs to runnable infrastructure. HarnessDev covers two stages. In Creation, the agent starts from a minimal seed and a small number of cases, then builds a complete execution system. In Evolution, it starts from its own created harness and iteratively revises it using downstream execution feedback, with the goal of improving benchmark performance. We then evaluate each constructed harness on capability (task success on held-out benchmarks) and efficiency (execution-token cost). The reported Creation results cover six creator LLMs, four domains, and five downstream benchmarks totaling 2,207 unique downstream instances, with hidden evaluation tasks withheld from development. We find that generated harnesses remain substantially behind mature human-engineered references on code and on search and research, while matching or exceeding the selected references on writing and machine-learning experimentation, with large variation in execution cost. Evolution produces some performance gains, but they are unstable and transfer only partially to held-out tasks. Experiments with a fixed runtime model further show that the gains depend strongly on the model executing the harness, indicating limited transfer across models.
Sources
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- DemoEvolve: Overcoming Sparse Feedback in Agentic Harness Evolution with Demonstrations
- From Failed Trajectories to Reliable LLM Agents: Diagnosing and Repairing Harness Flaws
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic Systems
- Evo-Bench: Can Language Models Improve Agent Harness?
- Towards Direct Evaluation of Harness Optimizers via Priority Ranking
- Continual Harness: Online Adaptation for Self-Improving Foundation Agents
- Recursive Harness Self-Improvement
- Meta-Harness: End-to-End Optimization of Model Harnesses
- OpenSage: Self-programming Agent Generation Engine
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- SEW: Self-Evolving Agentic Workflows for Automated Code Generation
- AgentBench: Evaluating LLMs as Agents
- The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?
- HarnessBank: Semantic Gene-Bank Search with Gated Verification for Agent-Harness Self-Evolution
- GAIA: a benchmark for General AI Assistants
- Code as Agent Harness
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties