Understanding Automated Program Repair Agents Through the Lens of Traceability: An Empirical Study
cs.SE, cs.AI
Submitted: 2025-06-10
Updated: 2026-08-29
Comments: Accepted for publication (ISSTA '26). Camera ready
Code: https://github.com/swe-bench/experiments
Project page: https://swe-agent-bench.github.io
License: http://creativecommons.org/licenses/by/4.0/
The gist: Automated Program Repair (APR) agents leverage large language models (LLMs) to autonomously diagnose and patch software bugs using planning, reasoning, and tools.
Terminology
Abstract
Automated Program Repair (APR) agents leverage large language models (LLMs) to autonomously diagnose and patch software bugs using planning, reasoning, and tools. Although these agents show strong performance on leaderboards such as SWE-bench, little is understood about how they take actions, where they fail, and how their behavior compares to human developers. In this paper, we present the first systematic analysis of these limitations using 5 state-of-the-art APR agents. We trace the full decision-making pipelines of the 5 APR agents across 500 real-world repair tasks, from issue description to patch validation. Our study reveals that, while agents excel at simple fixes, they struggle with logic-intensive bugs, often generating verbose, overfitted patches that pass existing test suites without solving the root cause. Test generation and regression test selection remain major bottlenecks, as agents fail to reproduce issues or run relevant regression tests. Moreover, many agents operate with primitive tooling (e.g. bash scripts) and do not have access to debuggers or program analysis tools. These findings highlight key limitations of current APR systems and motivate several directions for next-generation APR design, including but not limited to: (1) a shift-left approach emphasizing early, high-quality test generation and validation to reduce spurious fixes and improve semantic correctness; (2) richer, more integrated tool ecosystems; (3) diversified agent architectures that combine complementary strengths; and (4) benchmarks that prioritize semantic repair quality and test-generation fidelity over surface-level success metrics.
Sources
- Otter: Generating Tests from Issues to Validate SWE Patches
- TDD-Bench Verified: Can LLMs Generate Tests for Issues Before They Get Resolved?
- MASAI: Modular Architecture for Software-engineering AI Agents
- Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories
- Evaluating Software Development Agents: Patch Patterns, Code Quality, and Issue Complexity in Real-World GitHub Scenarios
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Are Coding Agents Generating Over-Mocked Tests? An Empirical Study
- S*: Test Time Scaling for Code Generation
- IRIS: LLM-Assisted Static Analysis for Detecting Security Vulnerabilities
- Comparing Human and LLM Generated Code: The Jury is Still Out!
- Process-Centric Analysis of Agentic Software Systems
- An Empirical Study on Failures in Automated Issue Solving
- Understanding Code Agent Behaviour: An Empirical Study of Success and Failure Trajectories
- Human-Agent versus Human Pull Requests: A Testing-Focused Characterization and Comparison
- Why Agentic-PRs Get Rejected: A Comparative Study of Coding Agents
- REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
- Promises, Perils, and (Timely) Heuristics for Mining Coding Agent Activity
- A Comprehensive Empirical Evaluation of Agent Frameworks on Code-centric Software Engineering Tasks
- Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties