Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

summary

Video file (mp4)

The gist

General-purpose agents like OpenClaw are increasingly used as autonomous tool users, but their coding ability remains difficult to measure under SWE-bench because generic agents do not inherently

In short

CLAW-SWE-BENCH is a new benchmark protocol designed to fairly compare different ways agents are built, called 'harnesses,' like OpenClaw. It standardizes evaluation by fixing everything except the harness itself, allowing researchers to see how much a specific agent design impacts performance under consistent rules for prompt and budget.

Key concepts

CLAW-SWE-BENCH
A multilingual benchmark protocol that treats agent harnesses as a controlled variable. It standardizes the evaluation stack—including prompt templates, execution settings, and patch extraction—to ensure that differences in results are due to the agent's design rather than inconsistent testing procedures.
Adapter Protocol
The first layer of CLAW-SWE-BENCH that creates a standardized interface between any agent harness and the benchmark. It requires each harness to implement specific methods like 'create_agent' and 'send_task,' allowing a shared orchestrator to control the entire lifecycle without needing to know the underlying code of the specific harness.
Fixed Base
The set of components in CLAW-SWE-BENCH that remain constant across all experiments, such as the prompt template, task set, execution container, and timeout. This fixed base ensures that any measured performance differences are attributable to variations in the agent harness or the underlying LLM being used.
Cost Accounting
A method for reporting agent evaluation results by including end-to-end run cost (in USD) alongside accuracy metrics like Pass@1. This allows researchers to plot accuracy and cost together on a single graph, treating cost as an important factor in evaluating coding agents.

Terminology used across episodes

This episode discusses

The paper

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks · Read on arXiv

TokenRhythm Technologies · Infinigence AI · City University of Hong Kong · SEE Fund · Peking University · Shanghai Jiaotong University · Beijing Jiaotong University · Tsinghua University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks".

Tom: General-purpose agents like OpenClaw are increasingly used as autonomous tool users,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, let's talk about the title of this paper, "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks," and who came up with it. Basically, they are focusing on creating a standardized way to test agent harnesses against coding tasks using a SWE-bench style framework.

Jane: The authors include quite a few researchers from places like Tsinghua University and Peking University, which suggests this work is coming from a strong academic environment, aiming for rigorous evaluation methods. They're essentially trying to solve the problem that existing reports conflate the model, the harness, and the task set into one score.

Lu: The core idea they are pushing is that because prior SWE-bench style reports package everything together—the prompt template, loop, timeout, extraction strategy—it becomes impossible to isolate what's causing a dip in performance when comparing different agent systems. This paper addresses that directly by isolating the harness dimension as the controlled variable.

Meng: From an engineering standpoint, that isolation is huge because it means if we see a performance difference between two agents, we can confidently point to the harness design instead of guessing which underlying LLM was slightly better or worse for that specific setup.

Lalam: I think this focus on the harness as an experimental variable is important because it shifts our focus from just optimizing the model's raw intelligence to designing a more effective agent framework that leverages that intelligence properly.

The paper's summary: Tom: Moving into the summary, what they are proposing is CLAW-SWE-BENCH H, which is a multilingual protocol for making these heterogeneous agent harnesses comparable under fair settings like prompt template and runtime budget. They are using a massive workload involving three hundred fifty real GitHub issue resolution tasks across eight languages and forty-three repositories.

Jane: The summary explains that they decompose the evaluation stack into a fixed base—things like the prompt template, execution container, patch extraction procedure, and evaluator—and then plug in a harness through an adapter protocol. This means any claw can run under this same evaluation protocol without needing to know how it's built internally.

Lu: The summary emphasizes that by fixing these base elements, they ensure that differences in metrics like Pass@one are attributable to the model or the specific harness dimensions being tested, not inconsistencies in how each system was originally evaluated.

Meng: That decomposition sounds like a very robust way to handle complexity; it suggests that we can build a standardized testing environment where we only need to swap out one piece—the harness—to see its effect clearly.

Lalam: It’s encouraging because it lays out exactly what needs to be standardized for meaningful comparison, giving us a clearer roadmap for designing better agent interaction layers in the future.

The paper's improvements: Tom: The paper highlights several specific improvements they've made to this evaluation setup, and one big improvement is treating cost-aware reporting as part of the benchmark design itself instead of just tacking it on at the end. They introduced a cost-aware, rank-aware selection method for their smaller subset.

Jane: That cost consideration is significant because it shows they aren't just interested in accuracy; they are concerned with how efficiently an agent can solve problems, which is a very practical concern for real-world applications where resource usage matters.

Lu: They also fixed several components that usually cause confusion, specifically standardizing the runtime and workspace by running tasks inside SWE-bench evaluation Docker images with the repository reset to the instance's base commit. This prevents agents from operating outside their intended context.

Meng: Fixing those contextual boundaries is crucial because if an agent can freely modify files or operate outside its defined scope, any performance metric we get becomes meaningless for deployment purposes, so that level of control is really important for me.

Lalam: The improvement in the patch and scoring contract is also key; they ensure that candidate solutions are collected from the repository state and the final prediction is computed against the base commit regardless of whether the harness outputs JSON or plain text. This standardization makes downstream evaluation much more reliable.

Conclusion: Tom: So, wrapping up, what we’ve seen here is that by introducing CLAW-SWE-BENCH H, they've created a system where agent harnesses can be rigorously compared on coding tasks while controlling for prompt structure and budget. The core finding is that the harness choice remains a first-order factor in performance.

Jane: They’ve shown that by treating the harness as a controlled experimental variable against fixed settings, we can interpret results on the same Pareto plane when accuracy and cost are plotted together, which gives us a much richer view of agent capability.

Lu: The implication for the research community is that this provides a standardized language and protocol for evaluating agents, moving away from system-specific reports toward a shared evaluation framework that allows for direct comparison of different agent designs.

Meng: From an engineering perspective, this gives us a tool to proactively design better agent loops; if we want an efficient solution, we can test the harness structure before even finalizing the LLM choice.

Lalam: This work really points toward a future where we can build AI systems not just around powerful models, but around intelligent infrastructure that is itself highly optimized for specific coding tasks.

Tom: And that brings us to an end for this paper discussion on "Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks." It’s a lot of foundational work setting up a much fairer way to measure agent effectiveness.

Jane: Indeed, it sets a high bar for how we should be evaluating autonomous agents going forward. It’s an important step in making agent development more systematic and transparent.

Lu: We're really excited about the potential this protocol opens up for testing novel agent architectures across different domains.

Meng: I just hope that when the community starts adopting this benchmark, it leads to more practical, efficient agents that we can actually deploy into production environments.

Lalam: I think this paper is a great step in ensuring that our AI systems are evaluated on their structural integrity as much as their functional output.

More episodes

← Home