Rethinking the Evaluation of Harness Evolution for Agents
cs.AI
Submitted: 2026-07-14
Updated: 2026-08-27
License: http://creativecommons.org/licenses/by/4.0/
The gist: We revisit the evaluation of automatic harness evolution for LLM agents.
Terminology
Abstract
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance on the same public benchmark. This protocol raises two fundamental concerns. First, harness evolution is itself an iterative search procedure that repeatedly evaluates and revises candidate harnesses using task feedback. As in agentic test-time scaling, it should therefore be compared with simple task-level search baselines under matched feedback and inference budgets to determine whether its gains arise from improved harness design or from additional search alone. Second, because the search and the final evaluation share the same benchmark, the reported gains risk overfitting to that specific task set. To address these concerns, we conduct an extensive evaluation comparing harness evolution with simple test-time scaling and discovery baselines under comparable feedback and inference budgets, and also evaluate evolved harnesses on held-out tasks to assess whether the discovered improvements generalize. Experiments on Terminal-Bench 2.1 with GPT-5.4 and Claude Opus 4.6 show that automatic harness evolution does not consistently outperform simple test-time scaling methods and exhibits limited generalization. Our results raise important questions about the effectiveness of automatic harness evolution and highlight the need for fairer evaluation protocols and benchmarks for automatic harness design. Our code is available at https://github.com/rethinking-harness-evolution.
Sources
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- Evaluating Test-Time Scaling of General LLM Agents
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Harnessing Agentic Evolution
- Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection