Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks".
Tom: General-purpose agents like OpenClaw are increasingly used as autonomous tool users,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, let's talk about the title of this paper, "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks," and who came up with it. Basically, they are focusing on creating a standardized way to test agent harnesses against coding tasks using a SWE-bench style framework.
Jane: The authors include quite a few researchers from places like Tsinghua University and Peking University, which suggests this work is coming from a strong academic environment, aiming for rigorous evaluation methods. They're essentially trying to solve the problem that existing reports conflate the model, the harness, and the task set into one score.
Lu: The core idea they are pushing is that because prior SWE-bench style reports package everything together—the prompt template, loop, timeout, extraction strategy—it becomes impossible to isolate what's causing a dip in performance when comparing different agent systems. This paper addresses that directly by isolating the harness dimension as the controlled variable.
Meng: From an engineering standpoint, that isolation is huge because it means if we see a performance difference between two agents, we can confidently point to the harness design instead of guessing which underlying LLM was slightly better or worse for that specific setup.
Lalam: I think this focus on the harness as an experimental variable is important because it shifts our focus from just optimizing the model's raw intelligence to designing a more effective agent framework that leverages that intelligence properly.
The paper's summary: Tom: Moving into the summary, what they are proposing is CLAW-SWE-BENCH H, which is a multilingual protocol for making these heterogeneous agent harnesses comparable under fair settings like prompt template and runtime budget. They are using a massive workload involving three hundred fifty real GitHub issue resolution tasks across eight languages and forty-three repositories.
Jane: The summary explains that they decompose the evaluation stack into a fixed base—things like the prompt template, execution container, patch extraction procedure, and evaluator—and then plug in a harness through an adapter protocol. This means any claw can run under this same evaluation protocol without needing to know how it's built internally.
Lu: The summary emphasizes that by fixing these base elements, they ensure that differences in metrics like Pass@one are attributable to the model or the specific harness dimensions being tested, not inconsistencies in how each system was originally evaluated.
Meng: That decomposition sounds like a very robust way to handle complexity; it suggests that we can build a standardized testing environment where we only need to swap out one piece—the harness—to see its effect clearly.
Lalam: It’s encouraging because it lays out exactly what needs to be standardized for meaningful comparison, giving us a clearer roadmap for designing better agent interaction layers in the future.
The paper's improvements: Tom: The paper highlights several specific improvements they've made to this evaluation setup, and one big improvement is treating cost-aware reporting as part of the benchmark design itself instead of just tacking it on at the end. They introduced a cost-aware, rank-aware selection method for their smaller subset.
Jane: That cost consideration is significant because it shows they aren't just interested in accuracy; they are concerned with how efficiently an agent can solve problems, which is a very practical concern for real-world applications where resource usage matters.
Lu: They also fixed several components that usually cause confusion, specifically standardizing the runtime and workspace by running tasks inside SWE-bench evaluation Docker images with the repository reset to the instance's base commit. This prevents agents from operating outside their intended context.
Meng: Fixing those contextual boundaries is crucial because if an agent can freely modify files or operate outside its defined scope, any performance metric we get becomes meaningless for deployment purposes, so that level of control is really important for me.
Lalam: The improvement in the patch and scoring contract is also key; they ensure that candidate solutions are collected from the repository state and the final prediction is computed against the base commit regardless of whether the harness outputs JSON or plain text. This standardization makes downstream evaluation much more reliable.
Conclusion: Tom: So, wrapping up, what we’ve seen here is that by introducing CLAW-SWE-BENCH H, they've created a system where agent harnesses can be rigorously compared on coding tasks while controlling for prompt structure and budget. The core finding is that the harness choice remains a first-order factor in performance.
Jane: They’ve shown that by treating the harness as a controlled experimental variable against fixed settings, we can interpret results on the same Pareto plane when accuracy and cost are plotted together, which gives us a much richer view of agent capability.
Lu: The implication for the research community is that this provides a standardized language and protocol for evaluating agents, moving away from system-specific reports toward a shared evaluation framework that allows for direct comparison of different agent designs.
Meng: From an engineering perspective, this gives us a tool to proactively design better agent loops; if we want an efficient solution, we can test the harness structure before even finalizing the LLM choice.
Lalam: This work really points toward a future where we can build AI systems not just around powerful models, but around intelligent infrastructure that is itself highly optimized for specific coding tasks.
Tom: And that brings us to an end for this paper discussion on "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks." It’s a lot of foundational work setting up a much fairer way to measure agent effectiveness.
Jane: Indeed, it sets a high bar for how we should be evaluating autonomous agents going forward. It’s an important step in making agent development more systematic and transparent.
Lu: We're really excited about the potential this protocol opens up for testing novel agent architectures across different domains.
Meng: I just hope that when the community starts adopting this benchmark, it leads to more practical, efficient agents that we can actually deploy into production environments.
Lalam: I think this paper is a great step in ensuring that our AI systems are evaluated on their structural integrity as much as their functional output.
TokenRhythm Technologies · Infinigence AI · City University of Hong Kong · SEE Fund · Peking University · Shanghai Jiaotong University · Beijing Jiaotong University · Tsinghua University
cs.LG, cs.CL
Submitted: 2026-06-10
Updated: 2026-09-28
Code: https://github.com/opensquilla/claw-swe-bench
Importance score: 90/100
The gist: General-purpose agents like OpenClaw are increasingly used as autonomous tool users, but their coding ability remains difficult to measure under SWE-bench because generic agents do not inherently
Key concepts
- CLAW-SWE-BENCH
- A multilingual benchmark protocol that treats agent harnesses as a controlled variable. It standardizes the evaluation stack—including prompt templates, execution settings, and patch extraction—to ensure that differences in results are due to the agent's design rather than inconsistent testing procedures.
- Adapter Protocol
- The first layer of CLAW-SWE-BENCH that creates a standardized interface between any agent harness and the benchmark. It requires each harness to implement specific methods like 'create_agent' and 'send_task,' allowing a shared orchestrator to control the entire lifecycle without needing to know the underlying code of the specific harness.
- Fixed Base
- The set of components in CLAW-SWE-BENCH that remain constant across all experiments, such as the prompt template, task set, execution container, and timeout. This fixed base ensures that any measured performance differences are attributable to variations in the agent harness or the underlying LLM being used.
- Cost Accounting
- A method for reporting agent evaluation results by including end-to-end run cost (in USD) alongside accuracy metrics like Pass@1. This allows researchers to plot accuracy and cost together on a single graph, treating cost as an important factor in evaluating coding agents.
Terminology
Summary
General-purpose agents like OpenClaw are increasingly used as autonomous tool users, but their coding ability remains difficult to measure under SWE-bench because generic agents do not inherently satisfy the required clean Docker workspace, patch, and prediction contract. This paper introduces CLAW-SWE-BENCH H, a multilingual benchmark protocol that makes heterogeneous agent harnesses comparable by treating the harness as a controlled experimental variable against fixed prompt, runtime budget, and evaluation settings.
The gist
CLAW-SWE-BENCH is a multilingual SWE-bench-style benchmark and adapter protocol that makes heterogeneous agent harnesses, or claws, comparable under fair settings including a fixed prompt, runtime budget, workspace contract, patch extraction procedure, and evaluator.
How it works
The core technical problem addressed is separating the evaluated LLM from the harness that turns the LLM into an agent. To achieve this comparability and attribution of results to harness design rather than inconsistent evaluation protocols, CLAW-SWE-BENCH decomposes the evaluation stack into a fixed base—prompt template, task set, execution container, per-instance timeout, patch extraction, and evaluator—plus a replaceable harness slot connected via a shared adapter protocol. This allows heterogeneous claws
to run under the same evaluation protocol.
The workload is constructed from two upstream SWE-bench-derived sources: SWE-bench-Multilingual (300 non-Python instances) and SWE-bench-Verified-Mini (50 human-validated Python instances), totaling 350 real GitHub issue resolution tasks across 8 languages and 43 repositories. All systems share the same outer budget, which includes a fixed base—prompt template, task set, execution container, per-instance timeout, patch extraction, and evaluator.
This shared control ensures that differences in Pass@1 are attributable to model or harness dimensions rather than inconsistent protocols.
Adapter Protocol and Execution Pipeline
The adapter protocol is the first layer of CLAW-SWE-BENCH. It standardizes the interface between a harness and the benchmark lifecycle by requiring each supported harness to implement abstract methods: create agent, send task, backup session, delete agent, and get docker args.
The shared orchestrator drives a run only through these methods without needing to know the underlying harness implementation.
At runtime, the shared orchestrator enforces an access boundary where container startup, repository reset to the instance’s base commit in /testbed, prompt instantiation from a fixed template (which forbids git add/commit), patch collection from repository state rather than agent final messages, and downstream SWE-bench evaluation are handled uniformly by the benchmark layer. This structure ensures that accuracy and cost can be interpreted in the same table and on the same Pareto plane.
Standardized Execution Pipeline Controls
CLAW-SWE-BENCH further fixes components that would otherwise confound harness comparisons:
-
Runtime and workspace: Each task runs inside its corresponding SWE-bench evaluation Docker image, with the repository reset to the instance’s base commit and mounted at /testbed. Future commits are cleaned to ensure agents only operate within the history boundary of the issue.
-
Prompt instantiation: Every instance is instantiated from a single task-prompt template that instructs the agent to work in /testbed and forbids modifying test files.
-
Patch and scoring contract: Candidate solutions are collected from repository state, and after termination, the runner computes the diff against the base commit to write a SWE-bench-compatible prediction, regardless of whether the harness natively produces JSON or plain text.
Benchmark Subsets and Cost Accounting
To lower the barrier to use, CLAW-SWE-BENCH Lite is released as an 80-instance low-cost subset. It uses a cost-aware, rank-aware selection method
over 17 calibration columns to optimize resolve-rate parity, pairwise ranking stability, and cost parity.
This subset reduces the full run cost to about 22.9% of the full benchmark while preserving key structural properties.
Evaluation metrics include Pass@1 (the fraction of instances whose submitted patch is marked RESOLVED by the SWE-bench evaluator), end-to-end run cost (TOTAL COST in USD), mean wall-clock duration, and cache hit rate. This multi-dimensional reporting allows accuracy and cost to be interpreted together on a Pareto plane, treating cost as first-class axis of SWE-style coding—agent evaluation.
Experimental Grids and Findings
The paper conducts two complementary studies: a model sweep (fixing OpenClaw and sweeping nine LLMs) and a claw sweep (fixing two models, GLM 5.1 and Qwen 3.6-flash, and sweeping five claws). These studies show that "harness choice is a first-order factor: under a fixed model, the claw spread reaches 12.5 pp on GLM 5.1 and 27.4 pp on Qwen 3.
Improvements for AI systems
As a fastidious researcher, I have analyzed the core contributions of the CLAW-SWE-BENCH paper. The primary innovation is moving beyond simple resolved rate
metrics to create a rigorous, controlled experimental variable for evaluating heterogeneous agent harnesses (claws) on real software engineering tasks, while explicitly accounting for resource costs and model variations.
Here are specific improvements to existing AI system evaluation and development workflows based on this research:
AI System Improvements Derived from CLAW-SWE-BENCH:
A shift from measuring agent performance
to measuring harness effectiveness
by treating the agent harness as a first-class controlled variable.
The implementation of a standardized, controllable evaluation pipeline that isolates the LLM's capability from the surrounding infrastructure (harness).
To enable precise, cost-aware comparison between different coding agents and their underlying models.
Improved AI System Capabilities:
An AI system can now be rigorously benchmarked not just on its final code quality (Pass@1), but on the efficiency of its execution—measuring the exact API cost, wall-clock duration, and cache hit rate required to achieve that result.
Developers can conduct harness-first
debugging: If an agent's performance drops, researchers can definitively determine if the issue lies in a flaw in the agent's internal loop (the claw design) or simply a weakness of the underlying LLM backbone.
System selection becomes data-driven and economically viable: Before deploying a new model or adopting a complex agent framework, engineers can use CLAW-SWE-BENCH Lite to predict how much cheaper (in terms of API spend and time) their chosen system will be on a full task set, allowing for cost/performance trade-off decisions that go beyond simple accuracy ranking.
The ability to identify harness brittleness: The research demonstrates that changing the agent loop (the claw) can cause performance drops comparable to changing the model itself, especially when using smaller models. This allows for proactive design choices regarding toolsets and stopping policies tailored to specific LLM capabilities.
A repeatable, transparent evaluation protocol is established: Any comparison between two agents or two models can be directly attributed to differences in harness design, task set distribution, and resource constraints, eliminating the ambiguity inherent in current leaderboard comparisons.
Sources
- Program Synthesis with Large Language Models
- Evaluating Large Language Models Trained on Code
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-Effi: Re-Evaluating Software AI Agent System Effectiveness Under Resource Constraints
- Measuring Coding Challenge Competence With APPS
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent Evaluation
- tinyBenchmarks: evaluating LLMs with fewer examples
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- SWE-smith: Scaling Data for Software Engineering Agents
- CodeClash: Benchmarking Goal-Oriented Software Engineering
- Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents
- AutoCodeRover: Autonomous Program Improvement
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks