HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, Yuan (Emily) Xue
Scale AI
cs.AI, cs.CL, cs.LG
Submitted: 2026-08-06
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 77/100
The gist: HarnessOpt-Bench: Evaluating LLMs at Harness Optimization is a paper by Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S.
Terminology
Summary
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization is a paper by Varun Ursekar, Apaar Shanker, Yash Maurya, Shehab Yasser, Vijay S. Kalmath, Veronica Chatrath, and Yuan (Emily) Xue from Scale AI (arXiv:2608.06301v1 [cs.AI], 6 Aug 2026). The paper introduces a benchmark for measuring how well frontier LLMs perform at harness optimization — the iterative and evaluation-guided improvement of a harness by an AI system
— where a harness is the prompts, tools, control flow, memory, and orchestration code surrounding
an LLM.
The authors motivate the work by noting that as LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness,
and that the same model can exhibit substantially different capabilities under different harnesses.
Because evaluation of stochastic agents is expensive and noisy,
harness optimization is a test of more than coding ability
: An optimizer must therefore diagnose failures from incomplete evidence, implement system-level changes, spend a limited evaluation budget, separate real improvement from noise, and decide what to deploy.
The paper makes four contributions:
-
A controlled benchmark for harness optimization: "We introduce H ARNESS O PT-B ENCH, in which an optimizer—an LLM paired with a coding harness—receives a target agent's seed harness, graded development and validation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search."
-
A trusted, reproducible evaluation protocol: "The suite comprises 4 downstream tasks with pinned seeds, fixed and non-overlapping development, validation, and test splits, graded disclosure, and recorded baselines for both the seed and off-the-shelf harnesses. A trusted execution environment, building on V E RO, enforces access and target-evaluation budgets, isolates held-out state, meters resource use, and versions every candidate for audit."
-
A controlled evaluation of optimizer models and coding harnesses: "We evaluate five frontier optimizer models under a shared coding harness and their respective native harnesses across 4 downstream tasks, yielding 111 scored optimizer runs, and evaluate two additional coding harnesses on one task."
-
Evidence of remaining capability gaps: "Release-level experiments show that H ARNESS O PT-B ENCH resolves variation among successive optimizer-model releases on OfficeQA. Instrumented search trajectories show that broader intervention is associated with greater held-out gain, whereas detailed failure-trace inspection is rarely used and is not positively associated with gain."
The paper formalizes harness optimization as a constrained, stochastic program-optimization problem.
An optimizer receives H0, interacts with Fθ under disclosure π and budget B, produces candidates H1,..., HT ∈ H, and nominates a final candidate H+.
The objective is to maximize the expected improvement of its nominated candidate over the pinned seed on the held-out partition
:
max [Eθ(H) − Eθ(H0)]
H ∈ H s.t. Σj cj ≤ B
Because raw improvement may not be comparable across tasks whose scoring scales differ,
the paper uses normalized gain:
g = [Eθ(H+) − Eθ(H0)] / [1 − Eθ(H0)]
where negative values indicate a nominated candidate worse than the seed.
-
Candidates and invariants:
A candidate harness H is an executable codebase. No semantic partition is imposed between prompts, tool definitions, memory, and control flow.
The optimizermay change H but not θ
(where θ = M, E, V fixes the models, environment, and verifier). -
Evaluation and disclosure: Cases are partitioned into
disjoint development, validation, and test sets, Ddev, Dval, and Dtest.
Development "reveals case inputs, per-case outcomes, and traces to support diagnosis; validation reveals only an aggregate score to support selection. The test partition is inaccessible during search and is evaluated by the trusted server only after the optimizer nominates a candidate." -
Budget:
the primary components of B are caps of 100 evaluation calls per partition and four full case passes on each of the development and validation partitions, plus a cap on total expendable target-model tokens.
-
Trusted execution: "Each optimizer is run in an isolated sandbox with evaluation results, task data, and the target harness in the filesystem. Similarly, each target agent rollout is performed in an ephemeral sandbox to control for noise introduced by environmental drift."
The suite comprises 4 downstream tasks, each with a small, deliberately untuned Python harness that leaves obvious headroom
:
-
OfficeQA (target model deepseek-v4-flash, split 49/98/99): a ∼130-line agent built on the OpenAI API with
three tools, a 24-turn loop, and a generic system prompt.
Baseline held-out score 0.341 ± 0.023. -
BrowseComp-Plus (deepseek-v4-flash, split 33/66/66): baseline 0.462 ± 0.020.
-
Terminal-Bench (grok-build, split 17/36/36): baseline 0.241 ± 0.009.
-
GAIA (gpt-5.4-mini, split 33/66/66):
The GAIA seed is a non-functional stub
witha measured-zero baseline,
so gain there is the raw held-out score,measuring building a working agent rather than improving a competent one.
Three of the four are competent but naive; the OfficeQA seed, for example, is a ∼130-line agent built on the OpenAI API, with three tools, a 24-turn loop, and a generic system prompt. GAIA's is a non-functional stub.
The seeds perform comparably or worse than the weakest
off-the-shelf harnesses, providing ample headroom for improvement.
Across all tasks, M = 1 (one pinned target model per task). Each task's baseline Eθ(H0) is measured once, averaged over K=3 independent rounds, and pinned for reproducibility; a nominated candidate is likewise scored three times per test case and averaged.
The authors evaluate five frontier optimizer models from three developers: claude-opus-5, claude-sonnet-5, gpt-5.6-sol, gpt-5.6-terra, and kimi-k3. Each model is paired with two coding harnesses: a shared harness (opencode), held fixed across all models, and the model's native harness (claude-code for Claude models, codex for GPT models, kimi-cli for Kimi). This gives 10 configurations in our core grid.
Two additional harnesses (goose and mini-swe-agent) are run on GAIA alone, and for two models (claude-opus and gpt-5.x), earlier releases are run on OfficeQA supporting the capability ladder.
-
Each held-out evaluation is
the mean of three attempts per test case
matching the K=3 pooling behind each task's pinned baseline. -
Each optimizer configuration is run twice; ranges are reported.
-
The optimizer model's inference is
metered but left uncapped, so these results estimate what is achievable when the optimizer model's reasoning is not the scarce resource.
-
Measurement resolution:
We estimate evaluation noise by scoring the same candidate twice on the same cases and carry the median discrepancy to the K=3 normalized-gain scale used for held-out scoring.
Task-specific resolution bands are: OfficeQA ±0.045, BrowseComp-Plus ±0.066, Terminal-Bench ±0.054, GAIA ±0.035.differences smaller than the band are treated as unresolved.
-
Composite model score (LSS-λ): Normalized gain is decomposed as
ḡmt = μ + τt + λm + εmtwhere τt absorbs task differences and λm is the optimizer model effect.LSSλ(m) = λ̂m
isthe model's task-adjusted mean performance relative to the evaluated models' grand mean, expressed in normalized-gain units; higher is better.
Model choice has a larger effect than coding-harness choice. Holding task and harness fixed, changing the optimizer model moves gain by 0.142 on average; holding task and model fixed, changing the harness moves it by 0.079. The model contrast is therefore about 1.8× larger.
The extremes separate more clearly than the middle. "The strongest configuration captures roughly two thirds of the available OfficeQA headroom and half of the BrowseComp-Plus headroom, while the weakest is unresolved from zero on BrowseComp-Plus and Terminal-Bench. Differences among intermediate configurations are often smaller than their round-to-round variation, supporting tiers rather than a complete ranking."
The LSS-λ model effects (with resolution ±0.058) place claude-opus-5 in tier 1 at +0.228, claude-sonnet-5 (+0.029), kimi-k3 (−0.014), and gpt-5.6-sol (−0.069) in tier 2, and gpt-5.6-terra (−0.174) in tier 3, where a shared tier means 'closer together than re-running the grid moves them' and ranks within a tier are not resolved.
Using two OfficeQA release series with the target model, seed, budget, and coding harness fixed within each series
: "Across 5 GPT releases, gain rises monotonically from +0.03 to +0.49, with three of its four steps exceeding the task's resolution band. Across 5 Claude Opus releases, gain ranges from +0.37 to +0.59; gain is non-monotonic, but the first-to-last spread exceeds the task's resolution band. This demonstrates that
H ARNESS O PT-B ENCH resolves variation among successive optimizer-model releases."
Broader search is associated with greater gain. The authors identified eight harness levers on OfficeQA before examining the other tasks: prompt, context management, step cap, retry and timeout policy, tool schema, answer extraction, retrieval policy, and reasoning effort.
The fraction touched during search is positively associated with gain on every task, with ρ ranging from +0.34 to +0.88.
This is exploration rather than final candidate breadth: An optimizer may inspect or modify a lever without retaining the change: one configuration touched three quarters of the levers and made seven edits but shipped the original seed.
Trace reading is not associated with higher gain. "The share of actions spent reading evaluation output is negatively associated with gain across the four tasks, from −0.31 to −0.64. Optimizers rely mainly on per-case score summaries: detailed trace spans were requested only 16 times by 7 of the 111 cells."
Case passes, not evaluation calls, bind. The median optimizer uses 8 calls (4%) but 82% of its case allowance, and 55 of 100 cells exhaust at least one partition's case budget. Thus the case allowance constrains search, whereas the call cap does not.
Visible validation scores are optimistic. "Most cells fall below the identity line: the submitted candidate's test score is lower than the best validation score observed during search. The held-out partition is therefore necessary to measure realized gain... we claim only that the visible best score is optimistic."
"Across the 20 model–task pairs evaluated under both conditions, the shared harness wins 11, the native harness wins 9, and 0 are tied. Native harnesses therefore have no consistent advantage. Restricting evaluation to a model's native tooling would not provide a reliable estimate of its harness-optimization ability."
However, In 11 pairs, the difference exceeds the task's resolution band, but the direction varies by model and task. The aggregate near-tie therefore reflects heterogeneous effects rather than uniformly small ones.
On GAIA — the only task whose harness axis has more than two levels
— "The harnesses do not hold their rank across models, and the native-harness advantage is concentrated rather than absent: both GPT models are four to five resolution bands better under codex, while the two Claudes and Kimi sit within a band or two of zero."
The paper concludes: "Harness engineering is becoming a model capability, not merely infrastructure around one. H ARNESS O PT-B ENCH makes that capability an empirical object: can a model diagnose, modify, and improve an agent as measured via a held-out reward? Current frontier models can, but unevenly. The strongest search broadly, yet their gains remain task-dependent and often too close to support a fine-grained ranking. By releasing H ARNESS O PT-B ENCH, we make this capability reproducible to measure and concrete to optimize. The next frontier is not merely better agents, but models that reliably make agents better."
The paper acknowledges several limitations:
-
"H ARNESS O PT-B ENCH is designed to be hack-resistant, not hackproof. The optimizer cannot access the test partition or alter the target model, environment, or verifier, but repeated development and validation feedback may still reward strategies specific to a fixed evaluation."
-
"The seed harness is itself a task-specific prior. Improving a mature agent tests diagnosis and refinement; starting from a stub tests construction. Our suite contains both regimes but does not vary seed complexity systematically."
-
Candidates are restricted to Python and each task uses one pinned target model. Generalization to other languages, runtimes, agent architectures, and target models remains untested.
The paper notes a dual-use tension, since the same capability that repairs and strengthens a benign agent could in principle be turned toward a harmful one.
It mitigates this because "the optimization target is always a bounded, benign downstream task (document QA, deep research, multi-step reasoning, or terminal use), the score is the task's own verifier on a held-out split and nothing else, and every optimizer runs inside a trusted harness that meters spend, versions every change for audit, and confines the target to a fixed model behind an allow-list."
-
OfficeQA (band ±0.045): claude-opus-5/opencode leads at 0.63; claude-opus-5/claude-code 0.59; claude-sonnet-5/claude-code 0.53; claude-sonnet-5/opencode 0.51; gpt-5.6-sol/codex 0.49; kimi-k3/kimi-cli 0.59; kimi-k3/opencode 0.41; gpt-5.6-sol/opencode 0.29; gpt-5.6-terra/opencode 0.14; gpt-5.6-terra/codex 0.07.
-
BrowseComp-Plus (band ±0.066): claude-opus-5/opencode leads at 0.48; claude-opus-5/claude-code 0.41; kimi-k3/kimi-cli 0.23; kimi-k3/opencode 0.16; claude-sonnet-5/opencode 0.15; gpt-5.6-sol/opencode 0.09; claude-sonnet-5/claude-code 0.07; gpt-5.6-sol/codex 0.03; gpt-5.6-terra/opencode 0.02; gpt-5.6-terra/codex −0.03.
-
Terminal-Bench (band ±0.054): claude-opus-5/opencode leads at 0.29; claude-opus-5/claude-code 0.18; kimi-k3/kimi-cli 0.16; claude-sonnet-5/opencode 0.15; gpt-5.6-sol/opencode 0.13; gpt-5.6-sol/codex 0.12; kimi-k3/opencode 0.12; claude-sonnet-5/claude-code 0.10; gpt-5.6-terra/opencode 0.04; gpt-5.6-terra/codex 0.01.
-
GAIA (band ±0.035, raw scores): gpt-5.6-sol/codex leads at 0.49; claude-opus-5/opencode 0.47; claude-opus-5/claude-code 0.42; claude-sonnet-5/claude-code 0.33; gpt-5.6-sol/opencode 0.31; kimi-k3/kimi-cli 0.31; gpt-5.6-terra/codex 0.30; kimi-k3/opencode 0.28; claude-sonnet-5/opencode 0.25; gpt-5.6-terra/opencode 0.17.
Improvements for AI systems
Improvement: Build self-improving agent systems around an explicit lever space
— prompt, context management, step cap, retry/timeout policy, tool schema, answer extraction, retrieval policy, reasoning effort — with tracking of which levers have been touched, modified, retained, or reverted. The paper found the fraction of levers touched is positively associated with held-out gain on every task (ρ = +0.34 to +0.88).
What the improved system can do: Systematically explore a wide intervention space instead of fixing the first failure it sees, achieving larger verified gains on held-out tests (e.g., capturing 2⁄3 of available OfficeQA headroom instead of shipping narrow patches).
Improvement: Rebalance the search loop around the actual binding constraint. The paper shows optimizers use only 4% of the call cap but 82% of the case allowance, with 55 of 100 cells exhausting at least one case budget. Design the loop to batch candidate evaluations per case pass, prioritize high-information cases, and reserve case passes for final confirmation.
What the improved system can do: Run many more candidate edits within a fixed evaluation budget without violating limits, escaping the case-pass bound
that currently halts search and limiting premature termination of optimization.
Improvement: Automatically generate actionable diagnostics — failure clusters by likely common cause, ranked hypotheses with confidence, per-case score distributions, estimated headroom per lever — rather than exposing raw traces. The paper found trace reading was rarely used (16 requests across 111 runs) and negatively associated with gain (−0.31 to −0.64), while per-case score summaries drove effective search.
What the improved system can do: Make high-signal decisions from compact evidence, avoiding the cost and failure mode of detailed trace inspection, and improve gain for models that currently stall because they won't or can't extract signal from traces.
Improvement: Embed task-specific measurement-resolution bands (e.g., ±0.045 OfficeQA, ±0.066 BrowseComp-Plus, ±0.054 Terminal-Bench, ±0.035 GAIA) into the decision loop: treat changes smaller than the band as unresolved, require repeated confirmations before retention, and report results with bands.
What the improved system can do: Stop chasing evaluation noise, avoid deploying false-positive improvements, and produce reproducible gains — with honest unresolved
outcomes instead of overclaimed rankings.
Improvement: Structure the system's own improvement loop with graded disclosure: a development partition with full feedback for diagnosis, a validation partition with aggregate scores only for selection, and a test partition that is inaccessible until a final candidate is nominated. Treat visible validation scores as optimistically biased (as the paper demonstrates).
What the improved system can do: Resist overfitting to its own validation signal and deliver gains that persist on held-out data — closing the systematic gap between best-visible validation scores and realized test scores.
Improvement: Train models specifically on the harness-optimization task — diagnosing failures from incomplete evidence, implementing system-level changes, managing a limited evaluation budget, separating real improvement from noise, and deciding what to deploy — using the benchmark's task-adjusted composite score (LSS-λ) as a training signal. The paper shows model choice moves gain 1.8× more than harness choice.
What the improved system can do: A model trained this way improves any agent it is paired with regardless of coding harness, providing the largest single lever on optimizer performance available in the stack.
Improvement: Detect whether the starting harness is a competent-but-naive agent or a non-functional stub, and switch strategy accordingly: construction-first (tools, control flow, verifier wiring) for stubs like GAIA; diagnosis-and-refinement-first for competent seeds like OfficeQA.
What the improved system can do: Achieve meaningful gains in both regimes — building a working agent from zero where baseline is zero (GAIA raw scores up to 0.49) and refining competent agents where headroom is partial (OfficeQA gains up to 0.63) — instead of applying one rigid search strategy.
Improvement: Decouple optimizer-model choice from harness choice: evaluate and select the coding harness per task and per model rather than defaulting to the model's native tooling. The paper found no consistent native-harness advantage (11 shared wins vs. 9 native wins across 20 model–task pairs) and heterogeneous direction (e.g., GPT models 4–5 resolution bands better under codex on GAIA, while Claude and Kimi show near-zero native advantage).
What the improved system can do: Gain multiple resolution bands of performance on tasks where the native pairing is suboptimal, and produce reliable cross-model comparisons instead of vendor-confounded rankings.
Improvement: Run every optimizer candidate in an isolated, versioned sandbox with metered target-model tokens, enforced evaluation budgets, pinned seeds, and a complete audit trail of every edit — the trusted-execution design of the benchmark applied to production systems.
What the improved system can do: Safely modify its own harness in production with full auditability, budget control, and reproducibility (e.g., K=3 pooled scoring), mitigating both operational risk and the dual-use concerns of models that can strengthen agents.
Improvement: Adopt a fixed task (target model, seed, budget, harness all pinned) as a continuous capability-ladder test for successive model releases, with release-to-release gain required to exceed the resolution band before progress is declared. The paper's OfficeQA ladder resolves monotonic gains (+0.03 → +0.49 across 5 GPT releases) and detects non-monotonicity in another series.
What the improved system can do: Give organizations a principled gate for whether a new model release genuinely improves harness-optimization ability — preventing deployment of releases whose gains are within noise and catching regressions early.
Abstract
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automated harness optimization -- the iterative and evaluation-guided improvement of a harness by an AI system -- both an important route to improving AI systems and a demanding capability for AI systems themselves. Yet the community lacks a common protocol for measuring how well frontier LLMs perform at this task. We introduce HarnessOpt-Bench, a benchmark for end-to-end harness optimization under expensive and stochastic evaluation. An optimizer, an LLM paired with a coding harness, receives a target agent's seed harness, graded evaluation feedback, and a fixed target-evaluation budget. It edits the harness and nominates a final candidate, which is scored by its normalized gain over the seed on a held-out test partition that remains inaccessible throughout search. A trusted execution environment enforces the evaluation boundary, meters target-agent resource use, and preserves candidate versions for audit. We evaluate 5 frontier LLMs as optimizers both under a shared coding harness and under their native harnesses across 4 downstream tasks, over 111 scored runs. Experiment results show that optimizer models separate more than the coding harnesses they act through, native harnesses are not consistently superior, and gains vary substantially across tasks and seed regimes. These results establish harness optimization as a measurable and discriminative capability with large space for improvement.
Sources
- Program Synthesis with Large Language Models
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Evaluating Large Language Models Trained on Code
- HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry
- Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines
- ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution
- Meta-Harness: End-to-End Optimization of Model Harnesses
- Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses
- RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- OfficeQA Pro: An Enterprise Benchmark for End-to-End Grounded Reasoning
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- A Self-Improving Coding Agent
- RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
- Harness-Bench: Measuring Harness Effects across Models in Realistic Agent Workflows
- Meta Context Engineering via Agentic Skill Evolution
- LLM-AutoDiff: Auto-Differentiate Any LLM Workflow
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection