Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents
cs.AI, cs.SE
Submitted: 2026-09-22
Updated: 2026-09-24
Comments: 16 pages, 6 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context.
Terminology
Abstract
Large language model (LLM) agents often handle streams of related tasks, yet standard harnesses repeatedly ask the model to reconstruct the same control decisions inside each task's context. We study whether task feedback can instead turn recurring control into reusable executable code, while reserving LLM calls for task-specific semantic reasoning. We introduce Growing Harness, a failure-guided training paradigm that learns the agent harness itself from a strategy-free scaffold that exposes fixed model and tool interfaces but encodes no task-solving controller. Function-level execution traces localize each failure to a bounded code surface, an optimizer repairs a window of failures jointly, and a success-first held-out gate rolls back repair sequences that harm prior capability. Accepted edits accumulate in one shared harness, allowing its control structure to emerge from task feedback. Across BrowseComp-Plus and WebArena-Verified with three deployment models from 4B to 120B parameters, Growing Harness achieves the highest mean success in five of six benchmark-model settings and trails the best mean by 0.7 pp. in the sixth. Relative to a Tool-Calling agent, it reduces LLM calls by 76.0-91.8% and deployed-agent inference cost by 74.4-98.6%. On WebArena-Verified, its success remains 44.7-45.3% across model scales, whereas Tool-Calling falls to 6.7% with the 4B model. Ablations show that trace-local edits, joint repair, and gate-based rollback each improve final success. These results show that persistent program growth can move recurring control out of model context and into low-cost code, yielding reusable specialist agents that remain effective with smaller deployment models.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Large Language Models as Tool Makers
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Automated Design of Agentic Systems
- MemoHarness: Agent Harnesses That Learn from Experience
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Meta-Harness: End-to-End Optimization of Model Harnesses
- AutoHarness: improving LLM agents by automatically synthesizing a code harness
- RouteLLM: Learning to Route LLMs with Preference Data
- gpt-oss-120b & gpt-oss-20b Model Card
- Optimizing Instructions and Demonstrations for Multi-Stage Language Model Programs
- Evolving Agents in the Dark: Retrospective Harness Optimization via Self-Preference
- Automatic Prompt Optimization with "Gradient Descent" and Beam Search
- VeRO: A Harness for Agents to Optimize Agents
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Do Agent Optimizers Compound? A Continual-Learning Evaluation on Terminal-Bench 2.0
- Large Language Models as Optimizers
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection