Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents".
Jane: Detailed Research Summary: Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents This research introduces Mid-Harness,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re diving into this paper today, "Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents." It sounds like they're tackling a really fundamental problem in how we make these AI agents work reliably. Basically, they’ve put a mechanism in place right at the boundary between the model generating actions and the harness that tests them.
Jane: Exactly, Tom. The title suggests it’s about scaling those actions by putting a verification step in there to make sure what the agent *does* actually helps it reach its goal consistently. It's not just about making better ideas; it's about making sure the actions chosen are actually good commands for the environment.
Lu: From my perspective, this paper is interesting because it moves beyond just improving reasoning within a single model run; they are systematically studying how adding computation at that specific juncture—the action generation point—impacts overall task success rates.
Meng: I’m curious about the practical side of this. If we're talking about terminal agents in engineering or science, what does "scaling actions" actually mean for us in terms of deployment? Is this just theoretical fluff, or is it something that translates into faster development cycles?
Lalam: From my perspective as an AI, I see this paper as highlighting how structure in the generation process leads to more robust outputs. It shows that the way we sample and check candidates matters more than just how much raw text a model spits out.
The paper's summary: Tom: Let's get into what they actually found in "Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents." They set up this framework where the model generates several potential actions, and then these candidates go through a verification step before one is actually sent to the environment.
Jane: So, the core idea is that simply generating many candidate actions isn't enough if you don't check them first. The paper shows that when you use a weak verification method, sampling more actions doesn't really help your success rate much at all.
Lu: They found a critical dependency: the benefit of sampling candidates is entirely governed by the quality of that verifier. If the verifier is weak, you don't gain much by generating more options from the model itself.
Meng: That makes sense from an engineering standpoint; if we have a buggy verification layer, spending compute on generating ten slightly different bad actions isn't worth the effort. So, what kind of verification mechanism did they test that actually showed promise?
Lalam: They found that among the mechanisms they tested, pairwise verification performed the best for this task. This suggests comparing candidates against each other gives a much clearer signal about which action path is viable.
The paper's improvements: Tom: Now, let's look at the specific improvements they suggest for this system. They highlight that using pairwise verification is the most effective way to identify good actions, and they also discuss how you can improve things by distilling knowledge from a stronger verifier back into the original model.
Jane: That distillation part is really interesting because it means you can get better performance without having to fundamentally change the action generator itself, which saves a lot of work for developers. It’s like giving the model a smarter filter after it’s already done its initial job.
Lu: The paper points out that command semantics and execution feasibility are persistent sources of disagreement between the distilled verifier and the teacher model, which accounts for about sixty-seven point four percent of verification failures in their offline analysis.
Meng: So, even when you distill knowledge, you still have to deal with those core issues of whether a command makes sense and whether it can actually be executed in the real system. That’s a concrete hurdle for us to consider during implementation.
Lalam: I think that sixty-seven point four percent figure is significant because it points directly to where we need our verifier training or refinement efforts to focus if we want real gains, especially concerning command semantics and execution feasibility.
Conclusion: Tom: So, wrapping things up on "Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents," the main conclusion is that action scaling is a promising area for using test-time compute in these agents because it shows tangible benefits when combined with proper verification.
Jane: They essentially proved that this systematic study lets us know exactly when adding sampling and verification translates into better trajectory success, especially when you use methods like pairwise comparison.
Lu: It’s a solid finding that gives us a clear roadmap on how to approach scaling for terminal agents by focusing compute where it yields the most benefit.
Meng: For practical application, this suggests that targeted refinement of actions via verification is more computationally efficient than just brute-forcing a huge number of trajectories, which is good news for our resource management.
Lalam: I think the biggest implication for us all is that we need to focus on building verifiers that are specifically tuned to understand command semantics and feasibility because those are the persistent failure points they identified.
Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl
NVIDIA
cs.CL, cs.AI, cs.MA
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: Project page: https://byungkwanlee.github.io/MidHarness-page/
Code: https://github.com/laude-institute/terminal-bench
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: This research introduces Mid-Harness, a novel framework designed to systematically investigate how allocating test-time compute at the model-harness boundary can enhance the reliability of actions
Key concepts
- Mid-Harness Framework
- This framework modifies how a language model generates actions by first sampling multiple potential actions and then checking them using a separate verification step before any action is used. It tests the relationship between generating many options and ensuring those options are actually good.
- Pairwise Verification
- This is a verification method where the system compares two candidate actions against each other to determine which one is superior. The research showed that comparing candidates directly provides the most useful information for finding viable paths for the agent.
- Verifier Strength Dictates Benefit
- The success of generating many action candidates depends entirely on how good the verification step is. A weak verifier offers little gain from sampling, whereas a strong verifier allows the system to successfully exploit better alternatives generated by the model.
- Distillation
- This involves taking the refined outputs from a stronger verification layer and feeding them back into an original model. This process helps improve success rates without changing the original action generator, showing how knowledge transfer across different layers can boost performance.
Terminology
Summary
This research introduces Mid-Harness, a novel framework designed to systematically investigate how allocating test-time compute at the model-harness boundary can enhance the reliability of actions generated by terminal agents. The core problem addressed is that while stochastic model generations produce candidate actions, their execution success is not guaranteed; a poor command (e.g., an incorrect package installation) can derail subsequent progress, even when superior alternatives exist within the same generative model.
Mid-Harness operates by inserting a sampling and verification step directly into the model's action generation pipeline, without altering the underlying generator or harness architecture. The proposed mechanism modifies the standard model call wrapper as follows:
Model-call wrapper (Mid-Harness)
async def generate(messages):
- return await llm.generate(messages)
-
candidates = await llm.generate(messages, n=N)
-
return await verify(llm, messages, candidates)
This structure allows researchers to study the interplay between action candidate generation (sampling N actions) and subsequent verification before a single action is forwarded to the environment.
The study establishes a critical dependency: verification governs the benefit of action sampling. The effectiveness of generating multiple candidates is highly contingent upon the quality of the verifier employed.
-
Verifier Strength Dictates Benefit: When using a weak verification mechanism, wider action sampling yields negligible improvements in trajectory success. Conversely, a strong verifier enables substantially more successful trajectories by effectively exploiting useful alternatives from the same generator.
-
Optimal Verification Mechanism: Among various mechanisms evaluated—including Listwise, Pointwise, and Pairwise verification—Pairwise verification performs best. This suggests that comparing candidate actions against each other provides the most informative signal for identifying viable paths.
-
Leveraging Verifier Distillation: A significant performance boost is achieved by distilling the responses from a stronger verifier back into the original model (e.g., TMAX-9B). This distillation process improves Pass@1 success without requiring any modification to the action generator itself, demonstrating how knowledge transfer across verification layers can enhance agent performance.
-
Semantic Disagreements: Offline analysis revealed that command semantics and execution feasibility are persistent sources of disagreement between the distilled verifier and the teacher model, accounting for a substantial portion (67.4%) of verification failures.
Mid-Harness is not only effective in isolation but also demonstrates synergistic benefits when combined with other scaling techniques:
-
Synergy with Trajectory Scaling: Combining Mid-Harness with established trajectory scaling methods, such as Best-of-T (T=3) and Sequential Refine (R=1), leads to higher overall success rates. Crucially, this combined approach achieves superior performance at a lower estimated token cost compared to simply generating a larger number of trajectories alone. This suggests that targeted action refinement is more computationally efficient than brute-force trajectory expansion.
-
Cross-Model and Benchmark Robustness: The framework demonstrated improved performance across different underlying models, distinct benchmarks, and various harness configurations, indicating its generalizability as a powerful tool for improving terminal agent reliability.
The central conclusion of the paper is that action scaling represents a promising target for test-time compute scaling in terminal agents. Mid-Harness provides the necessary systematic study to determine precisely when additional computation (sampling and verification) translates into tangible improvements in trajectory success, proving its compatibility with both parallel and sequential trajectory scaling strategies.
Despite its strong findings, the authors acknowledge several limitations that point toward future research:
-
Gap to Frontier Verification: The distillation approach leaves a substantial gap when compared to
frontier verification
methods, motivating the need for better training algorithms, such as reinforcement learning specifically designed for verifiers. -
Lack of Gold Labels: The evaluation inherently lacks gold action labels, which limits the direct measurement of verification correctness and candidate coverage—a crucial area for future rigorous validation.
-
Failure Analysis: The study notes that command semantics and execution feasibility are primary drivers of failure in distilled-verifier scenarios, reinforcing the need for verifiers to be highly attuned to these aspects.
The provided supplementary text from Candidate B describes a different type of problem-solving context—a code refactoring/debugging task involving transitive symmetric aliases and lexicographical ordering. While this is not directly related to the terminal agent paper, it illustrates the type of concrete, fix-oriented work that contrasts with purely observational debugging.
Improvements for AI systems
Here are specific improvements for AI systems based on the findings in Mid-Harness: Scaling Actions Between Model and Harness,
detailing what these improved systems can achieve:
) Mid-Harness (Action Verification) Integration
The core improvement is integrating a dedicated action verification step at the model-harness boundary. This moves beyond simple generation to a structured process of Generate Candidates → Verify Candidates → Execute Best Candidate.
-
The AI system can be architected with a fixed Generator (e.g., TMAX-9B) and Harness, but an added
Mid-Harness
wrapper that samples multiple candidate actions from the generator based on the current history. -
Before execution, this wrapper passes these candidates to a Verifier (which can be zero-shot or distilled). The Verifier's role is not just to judge plausibility but to explicitly evaluate the candidate against the task requirements, considering correctness, progress toward the goal, and risk of state change.
-
The system then executes only the single action chosen by the Verifier.
) Improved System Capabilities (What they can do)
This Mid-Harness architecture enables several high-value capabilities:
-
Active Error Correction and Recovery: The system can actively correct its own mistakes in real-time. If a generated command is flawed (e.g., wrong package install or incorrect code edit), the Verifier identifies this failure and selects a candidate that resolves the specific blocker, even if other candidates seem plausible in isolation.
-
State-Aware Search Space Pruning: Instead of blindly exploring all possible next actions, the system can use verification to prune irrelevant branches of the search tree. This means it spends its limited compute only on actions that demonstrably advance toward the task completion goal or fix known errors.
-
Enhanced Reliability on Complex Tasks: For long-horizon tasks (like software engineering), this system dramatically increases success rates (e.g., Pass@1 improvements of up to 80% compared to the base agent). It can reliably navigate complex dependencies and configuration changes where simple stochastic generation fails, as demonstrated by its ability to solve tasks where the base agent succeeds in less than 5% of runs.
-
Cost-Effective Scaling: The system proves that allocating test-time compute (for verification) is more effective than simply generating more trajectories or increasing reasoning effort for a single candidate. This allows developers to achieve higher reliability at lower estimated token costs by optimizing the
action scaling axis.
) Advanced Verification Strategies (How to make it smarter)
To maximize performance, the following specialized verifier mechanisms should be employed:
-
Pairwise Comparison (The Gold Standard): The system should use a pairwise verifier mechanism. This involves comparing two candidate actions against each other to generate a preference score based on their likely impact on the task progress and state change, rather than scoring them in isolation.
-
Knowledge Distillation for Verifiers: The system can benefit from training the Verifier model using
distilled
knowledge from a powerful frontier model (like GPT-5.6 Sol). This allows the smaller, fixed generator to leverage the superior judgment of a much larger model without needing to serve that massive model at inference time, significantly narrowing the gap between weak and strong verification capabilities. -
Decision-Only Verification for Efficiency: For high-throughput scenarios, using
decision-only
verifiers (which output only a single A/B token instead of a full rationale) can reduce inference costs by 20–24% while maintaining substantial gains in Pass@1, offering an optimal trade-off between performance and latency.
) Specific Application Examples:
The improved AI system can be used to:
-
Automatically resolve complex debugging scenarios (e.g., fixing specific error codes via configuration changes rather than general code edits).
-
Successfully complete multi-step software development tasks by correctly sequencing dependency installations and environment setups.
-
Navigate ambiguity in ambiguous terminal environments where multiple valid commands exist, by using pairwise comparison to select the action that minimizes risk or maximizes progress toward a known sub-goal.
Abstract
Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.
Sources
- Tmax: A simple recipe for terminal agents
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Rethinking Thinking Tokens: LLMs as Improvement Operators
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Training Verifiers to Solve Math Word Problems
- Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents
- Scaling Test-time Compute for LLM Agents
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Scaling Test-Time Compute for Agentic Coding
- GenSelect: A Generative Approach to Best-of-N
- $V_1$: Unifying Generation and Self-Verification for Parallel Reasoners
- CRIU -- Checkpoint Restore in Userspace for computational simulations and scientific applications
- Meta-Harness: End-to-End Optimization of Model Harnesses
- On Data Engineering for Scaling LLM Terminal Capabilities
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- OpenThoughts-Agent: Data Recipes for Agentic Models
- Endless Terminals: Scaling RL Environments for Terminal Agents
- FrogNano: Training a 4B Coding Agent via Online Task Synthesis
- SWE-RM: Execution-free Feedback For Software Engineering Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering