Scalable Supervision for Software Agents via Patch Reasoning
cs.CL, cs.SE
Submitted: 2025-10-26
Updated: 2026-08-26
Comments: EMNLP'26 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: While language model agents have advanced software engineering, existing test-based supervision is limiting its scalability on real-world issues.
Terminology
Abstract
While language model agents have advanced software engineering, existing test-based supervision is limiting its scalability on real-world issues. The reason is twofold: (1) high-coverage tests are naturally rare in the wild, and (2) building and running test sandbox is heavy and fragile. To unlock supervision scaling, we propose R4P, a reasoning-based method that provides scaffold-agnostic rewards. R4P uses a group-wise training objective, enabling it to verify multiple patches against each other's modification and gain a dense reward for supervising agents without executing tests or relying on specific agent trajectories. R4P achieves 72.2% Acc. for verifying patches from SWE-bench, competitive with proprietary models. To show the downstream practical utility of R4P, we design and train an execution-free scaffold, Mini-SE, with pure RL via R4P. Mini-SE achieves 26.2% Pass@1, showing a 10.0% improvement over the original Qwen3-32B, and can be further improved to 32.8% with R4P for test-time scaling on patch selection. The stable scaling curves illustrate that though imperfect, R4P can still reliably support downstream tasks at scale.
Sources
- RM-R1: Reward Modeling as Reasoning
- Trae Agent: An LLM-based Agent for Software Engineering with Test-time Scaling
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Repo2Run: Automated Building Executable Environment for Code Repository at Scale
- R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement
- SoRFT: Issue Resolving with Subtask-oriented Reinforced Fine-Tuning
- An Empirical Study on LLM-based Agents for Automated Bug Fixing
- Training Software Engineering Agents and Verifiers with SWE-Gym
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Vulnerability Detection: From Formal Verification to Large Language Models and Hybrid Approaches: A Comprehensive Overview
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
- SWE-smith: Scaling Data for Software Engineering Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering