SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring
Yuling Shi, Jinghan Xu, Kelin Fu, Wenhao Zeng, Shilin He, Lei Zhang, Yue Liu, Zelin Zhao, Terry Yue Zhuo, Jialun Cao, Siyu Ye, Tianyu Liu, Kai Cai, Shing-Chi Cheung, Xiaodong Gu
Shanghai Jiao Tong University · Peking University · The Hong Kong University of Science and Technology · Douyin Group · University of Chinese Academy of Sciences · National University of Singapore · Monash University
cs.CL, cs.SE
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: Published as a conference paper at COLM 2026
Code: https://github.com/nasa/fprime
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: SWE-Bench ProMax is an expert-curated, multilingual code refactoring benchmark designed to address three critical gaps in current AI coding evaluation: the saturation of existing benchmarks, the lack
Terminology
Summary
SWE-Bench ProMax is an expert-curated, multilingual code refactoring benchmark designed to address three critical gaps in current AI coding evaluation: the saturation of existing benchmarks, the lack of evaluation quality assurance, and the absence of refactoring as a benchmark domain. It comprises 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust) and 70 repositories, selected from 29,782 initial candidates through a rigorous three-stage pipeline. Every instance undergoes expert curation: issue descriptions are rewritten from scratch to eliminate ambiguity and answer leakage, and test suites are manually reviewed to remove overly narrow tests that reject valid solutions and overly broad tests that check unstated requirements. Evaluation of six frontier models reveals that SWE-Bench ProMax remains far from saturated: the best model (GPT-5.2) achieves only 41.2% resolve rate, and no model dominates across all languages. Our analysis uncovers two key findings. First, open-weight models are highly competitive with proprietary ones on large-scale refactoring: GLM-5 and Qwen3.5 come within five points of the best proprietary model at a fraction of the cost, and manual trajectory inspection shows that higher spending often reflects unproductive exploration cycles rather than genuine progress. Second, the dominant failure mode is incomplete refactoring: agents consistently modify fewer files than the gold patch requires, indicating that sustained cross-file coordination remains a fundamental bottleneck for current AI coding agents.
The benchmark operates at a fundamentally different scale than existing benchmarks: 30% of instances modify more than 10 files and 32% require over 200 lines of code, compared to SWE-bench Verified where 86% of instances modify only a single file. Gold patches average 11.4 source files and 261.6 lines of code (8,179.5 tokens), with the most complex instances modifying up to 182 source files. Test patches add a further 4.5 files and 185.5 lines on average, bringing the total to 15.9 files per instance. Issue descriptions average 685.3 tokens, providing detailed specifications of the required refactoring.
The construction pipeline progresses from automated collection through environment validation to expert-driven curation. Stage 1 uses the GitHub API to identify candidate repositories with at least 500 stars, an approved open-source license, and a primary language among the seven target languages, extracting commits submitted after January 2025 whose messages contain refactor
but not bug fix
and that modify both test and non-test files. Stage 2 constructs isolated Docker environments containing the repository at the pre-refactoring state with all dependencies installed, applies the gold patch, and executes the full test suite, discarding instances where a working environment cannot be established or the gold patch fails. Stage 3 involves human experts working with LLM assistance to transform validated instances into high-quality benchmark tasks through four steps: commit analysis, quality filtering (removing tasks confined to a single file, those with too few lines of code, overly narrow tests that enforce specific implementation details, and overly broad tests that check behavior beyond scope), problem statement rewriting (producing precise, self-contained specifications verified to serve as both necessary and sufficient conditions for the gold patch), and human verification.
The evaluation uses two agent scaffolds: mini-swe-agent, a minimal reimplementation of the widely adopted SWE-agent scaffold, and OpenHands, an open platform for generalist software agents with a richer runtime. A step limit of 300 and a cost limit of 10 per instance are set under both scaffolds. Six frontier models are evaluated: proprietary models Gemini-3-Pro, Claude Sonnet 4.6, and GPT-5.2, and open-weight models GLM-5, Kimi-K2.5, and Qwen3.5.
Under OpenHands, GPT-5.2 achieves the best resolve rate at 41.2%, followed by Claude Sonnet 4.6 at 38.8%, GLM-5 and Qwen3.5 at 36.5%, and Kimi-K2.5 at 32.9%. Gemini-3-Pro achieves only 19.4%. Under mini-swe-agent, all models perform substantially worse, with GPT-5.2 dropping to 21.8%, suggesting that richer runtime tooling is particularly beneficial for large-scale refactoring tasks. Open-weight models are dramatically more cost-efficient: GLM-5 resolves 36.5% at just 0.24 per instance—roughly one-twentieth the cost of Claude Sonnet 4.6 (4.77)—and Kimi-K2.5 achieves 32.9% at 0.72.
Performance varies substantially across languages. Claude Sonnet 4.6 leads on TypeScript (53.6%) and Rust (63.6%), GLM-5 leads on Java (34.6%), GPT-5.2 performs best on Python (48.3%) and C (75.0%), Kimi-K2.5 achieves its best result on Go (43.5%), and Qwen3.5 performs best on C++ (54.5%). TypeScript and Rust show surprising variance: Claude Sonnet 4.6 achieves 53.6% on TypeScript and 63.6% on Rust, while Gemini-3-Pro achieves 0.0% on TypeScript and Kimi-K2.5 only 18.2% on Rust.
Agent behavior analysis reveals two complementary failure patterns. The dominant failure mode is incomplete refactoring: agents identify and modify some affected files but fail to propagate changes to all locations requiring coordinated updates. Both Claude Sonnet 4.6 and Kimi-K2.5 closely track the gold patch distribution for small changes (up to 5 files) but diverge sharply for larger patches—whereas the gold patch cumulative distribution reaches 90% only around 20 files, both agents reach 90% by approximately 10 files. This partial coverage means that even when the central logic is correctly refactored, the test suite fails because downstream dependencies remain inconsistent. The second pattern concerns interaction rounds: successful resolutions complete in markedly fewer interaction rounds, with pass curves rising steeply and plateauing early, while fail curves are shifted rightward and more gradual. Agents entering unproductive cycles—repeatedly reading files, attempting edits, encountering test failures, and reverting—consume many rounds without expanding the scope of their modifications.
Cost analysis shows that higher cost does not translate to proportionally higher resolve rate. Claude Sonnet 4.6 is the most expensive model (4.77 per instance, 117.9 steps on average) yet trails GPT-5.2 (41.2% at 3.60), and Gemini-3-Pro spends 1.49 per instance for only 19.4%. Qwen3.5 takes the most steps under both scaffolds (155.4 and 141.2 on average) yet never leads overall, and under mini-swe-agent achieves the lowest resolve rate (20.6%), suggesting that excessive exploration without effective cross-file coordination is counterproductive.
Task category analysis using multi-label classification reveals that refactoring-related categories dominate—Refactoring Cleanup appears in 66.5% of instances and API Interface Change in 65.3%—but a substantial fraction of instances simultaneously involve new features (43.5%), bug fixes (41.2%), or documentation updates (21.2%). 46.5% of instances span three or more categories, with no instance involving fewer than two. The co-occurrence matrix reveals strong coupling between API changes and refactoring cleanup (79 co-occurrences), as well as between refactoring and bug fixing (48 co-occurrences), reflecting how restructuring code frequently exposes latent defects that must be addressed concurrently. Required skills analysis shows near-universal prevalence of cross-file reasoning (99.4%) and API semantics (98.8%), with pattern matching (91.8%), interface contract reasoning (97.1%), data flow (88.8%), domain knowledge (79.4%), and type system reasoning (50.6%) also common.
The benchmark includes representative instances across all seven languages, such as refactoring NASA's F'Prime flight software framework requiring coordinated changes across 244 files, adding hour-level time resolution to PlantUML's Gantt engine across 94 files, renaming motor protocol configuration fields in Betaflight across 62 files, unifying scalar arguments across tensor operations in the Burn deep learning framework across 49 files, refactoring upload stream buffering across drivers in OpenList across 47 files, centralizing provider output-format handling in Google's langextract across 30 files, and unifying destroyOnHidden across Ant Design components across 27 files.
Improvements for AI systems
Improvements to AI systems:
-
Cross-file coordination module: Add a planning mechanism that explicitly enumerates all files affected by a refactoring before editing begins, with a checklist that tracks propagation across dependencies. This directly targets the dominant failure mode where agents modify fewer files than required.
-
Cost-aware exploration governor: Implement a dynamic step-allocation policy that detects unproductive cycles (repeated reads, failed edits, reverts) and forces a re-planning phase after N consecutive failed attempts, reducing wasted interaction rounds without sacrificing success rate.
-
Language-adaptive strategy selection: Build a router that selects different agent strategies per programming language based on observed performance patterns (e.g., TypeScript/Rust benefit from deeper type-aware analysis, Python/C benefit from faster iteration, Go/C++ benefit from intermediate exploration).
-
Multi-category task decomposition: For instances involving overlapping concerns (refactoring + bug fix + API change), train the agent to explicitly separate sub-tasks and solve them sequentially rather than interleaving, reducing cognitive load and partial coverage.
-
Patch-size-aware test feedback: Modify the test harness to provide intermediate feedback on how many files/lines have been successfully modified relative to expected scope, enabling agents to self-correct before final test execution.
-
Open-weight model optimization: Since open models (GLM-5, Qwen3.5) achieve near-proprietary performance at 1/20th cost, develop specialized fine-tuning on large-scale refactoring trajectories to close the remaining 5-point gap, making high-quality coding assistance accessible at scale.
What the improved AI system can do:
-
Consistently modify all required files in multi-file refactorings (e.g., 10+ file changes) with >80% completion rate, up from current 50% partial coverage.
-
Complete refactoring tasks in 30-40% fewer interaction rounds by avoiding unproductive exploration loops.
-
Adapt its approach per language: achieving >60% resolve rate on TypeScript/Rust via type-aware planning, >50% on Python/C via rapid iteration, and >45% on Go/C++ via balanced exploration.
-
Handle complex mixed tasks (refactoring + bug fix + API change) with 20% higher success rate by decomposing into sequential sub-problems.
-
Self-monitor progress against expected patch scope and request additional context when stuck, reducing silent failures.
-
Deliver proprietary-level performance (40%+ resolve rate on SWE-Bench ProMax) at under 0.50 per instance using open-weight models, enabling affordable large-scale code modernization.
Abstract
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a recent audit found that nearly 60% of unsolved SWE-bench Verified instances contain flawed tests -- either overly narrow tests that reject correct solutions or overly broad tests that check unstated requirements -- and that frontier models can verbatim reproduce gold patches from training data. Code refactoring, which requires coordinated, behavior-preserving changes across many files, offers a substantially harder and more realistic test of agent capability, yet remains underserved by current benchmarks. We introduce SWE-Bench ProMax, an expert-curated, multilingual code refactoring benchmark of 170 instances drawn from real commits across seven programming languages (Python, Java, TypeScript, Go, C, C++, and Rust). Every instance undergoes rigorous, multi-stage curation that directly addresses the quality problems identified in prior benchmarks: issue descriptions are rewritten from scratch to provide precise, unambiguous specifications, and test suites are manually reviewed to remove overly narrow and overly broad tests. Tasks with insufficient complexity or limited cross-file scope are filtered out, yielding a benchmark of challenging, large-scale refactoring tasks that average 11.4 modified files and 261.6 lines of code per instance, substantially exceeding the scale of existing benchmarks. Experiments with frontier models under two agent scaffolds show that the best model achieves only 41.2% resolve rate, confirming that SWE-Bench ProMax presents a meaningful and unsaturated challenge for current AI coding agents. Our benchmark is available at https://huggingface.co/datasets/swe-bench-promax/SWE-Bench-ProMax.
Sources
- Program Synthesis with Large Language Models
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- Qwen3-Coder-Next Technical Report
- Evaluating Large Language Models Trained on Code
- SWE-Exp: Experience-Driven Software Issue Resolution
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- SWE-Marathon: Can Agents Autonomously Complete Ultra-Long-Horizon Software Work?
- daVinci-Env: Open SWE Environment Synthesis at Scale
- SWE-MeM: Learning Adaptive Memory Management for Long-Horizon Coding Agents
- RefactorBench: Evaluating Stateful Reasoning in Language Agents Through Code
- SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
- Agentic Refactoring: An Empirical Study of AI Coding Agents
- DeepSWE: Measuring Frontier Coding Agents on Original, Long-Horizon Engineering Tasks
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Measuring AI Ability to Complete Long Software Tasks
- SWE-EVO: Benchmarking Coding Agents in Long-Horizon Software Evolution Scenarios
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- SWE-QA: Can Language Models Answer Repository-level Code Questions?
- SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering