BoostAPR: Boosting Automated Program Repair via Execution-Grounded Reinforcement Learning with Dual Reward Models
cs.AI, cs.SE
Submitted: 2026-05-09
Updated: 2026-09-18
Comments: 21 pages, 2 figures. Accepted at ICML 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs.
Terminology
Abstract
Reinforcement learning for program repair is hindered by sparse execution feedback and coarse sequence-level rewards that obscure which edits actually fix bugs. We present BoostAPR, a three-stage framework addressing these challenges: (1) supervised fine-tuning on execution-verified demonstrations with reasoning traces, (2) training dual reward models--a sequence-level assessor and a line-level credit allocator--from execution outcomes, and (3) PPO optimization where the line-level model redistributes rewards to critical edit regions. This line-level credit assignment operates at an intermediate granularity naturally suited to code changes. Trained on SWE-Gym and evaluated on four benchmarks, BoostAPR achieves 40.7% on SWE-bench Verified (+22.9pp over base model), 24.8% on Defects4J (Python-to-Java transfer), 84.5% on HumanEval-Java, and 95.0% on QuixBugs, achieving competitive results among open-source models with strong cross-language generalization.
Sources
- Dense Reward for Free in Reinforcement Learning from Human Feedback
- Evaluating Large Language Models Trained on Code
- RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning
- MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
- Qwen2.5-Coder Technical Report
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement
- Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
- VAnim: Rendering-Aware Sparse State Modeling for Structure-Preserving Vector Animation
- Proximal Policy Optimization Algorithms
- HybridFlow: A Flexible and Efficient RLHF Framework
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- Alleviating Attention Hacking in Discriminative Reward Modeling through Interaction Distillation
- SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
- Reward Auditor: Inference on Reward Modeling Suitability in Real-World Perturbed Scenarios
- AutoCodeRover: Autonomous Program Improvement
- Agentless: Demystifying LLM-based Software Engineering Agents
- SWE-Fixer: Training Open-Source LLMs for Effective and Efficient GitHub Issue Resolution
- From Reaction to Anticipation: Proactive Failure Recovery through Agentic Task Graph for Robotic Manipulation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection