Coding Agents Have Converged: Why the SWE-bench Leaderboard Can No Longer Order Its Top Entries, and What to Measure Instead
cs.SE, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: Accepted at ADMA 2026 (International Conference on Advanced Data Mining and Applications), Special Session on Responsible Data Intelligence. Camera-ready version, 15 pages, 4 figures
Code: https://github.com/Adkid-Zephyr/resolution-audit
License: http://creativecommons.org/licenses/by/4.0/
The gist: Small differences on coding-agent leaderboards are often read as an ordering of systems.
Terminology
Abstract
Small differences on coding-agent leaderboards are often read as an ordering of systems. We audit whether the published verdicts support this reading, using 254 SWE-bench submissions across four splits without running models. On Verified, the leading two entries each resolve 396 of 500 instances. The top ten share 285 successes and 51 failures, leaving 164 instances that distinguish their outcomes. Frontier solution sets have median nesting 0.935 against a score-implied baseline of 0.774, indicating strongly shared successes. Scores also depend on the evaluated model-scaffold pair: observed within-model scaffold ranges reach 29.8 percentage points, compared with the 8.8-point spread of the top thirty. Six of nine cell-mean interaction tests remain significant after Holm correction, although this observational design does not identify causal scaffold effects. Exact paired McNemar tests separate none of the 29 adjacent Verified top-thirty pairs at alpha=0.05, while the larger Test split separates 14 of 23. A stated leader-based rule yields three descriptive tiers, or two after Holm correction; non-rejection does not establish equivalence. We release the partition and a five-step audit protocol that profiles shared outcomes, tests paired differences, reports grouping sensitivity, and estimates the instance budget needed for resolution. The results motivate reporting comparison-set-specific resolution and model-scaffold provenance instead of interpreting small aggregate gaps as established rank differences.
Sources
- SWE-Bench+: Enhanced Coding Benchmark for LLMs
- Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
- Resolution Diagnostics for Paired LLM Evaluation
- Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation
- AgentLens: Revealing The Lucky Pass Problem in SWE-Agent Evaluation
- Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
- Automated Benchmark Auditing for AI Agents and Large Language Models
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
- Establishing Best Practices for Building Rigorous Agentic Benchmarks
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties