A 2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents
cs.CL, cs.SE
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted to EMNLP 2026 Main
Code: https://github.com/donian00/A2Agent
License: http://creativecommons.org/licenses/by/4.0/
The gist: Localizing issue-relevant code regions is a critical step in automated software engineering.
Terminology
Abstract
Localizing issue-relevant code regions is a critical step in automated software engineering. However, due to their reliance on sparse trajectory-level signals, existing methods cannot identify which per-turn actions are effective and often discover correct code regions during exploration but fail to commit them. To address these limitations, we propose an action-aware reinforcement learning method that combines a per-turn reward sequence rewarding both the discovery and commitment of gold code regions with an action-level advantage estimation scheme that isolates each action's credit by grouping turns sharing the same exploration context. Extensive evaluations show that our method improves the average F1 over the state-of-the-art (SOTA) by 1.58% on SWE-Bench Verified and 8.55% on SWE-Bench Pro, with our 4B model outperforming baselines up to 8x larger. Our code is available at https://github.com/donian00/A2Agent.
Sources
- Evaluating Large Language Models Trained on Code
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- Reinforced Self-Training (ReST) for Language Modeling
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization
- Hierarchy-of-Groups Policy Optimization for Long-Horizon Agentic Tasks
- LoRA: Low-Rank Adaptation of Large Language Models
- Issue Localization via LLM-Driven Iterative Code Graph Searching
- VinePPO: Refining Credit Assignment in RL Training of LLMs
- Tool-integrated Reinforcement Learning for Repo Deep Search
- AWAC: Accelerating Online Reinforcement Learning with Offline Datasets
- Training Software Engineering Agents and Verifiers with SWE-Gym
- Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
- SweRank: Software Issue Localization with Code Ranking
- Code Llama: Open Foundation Models for Code
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- CodeScout: An Effective Recipe for Reinforcement Learning of Code Search Agents
- Improving Code Localization with Repository Memory
- Agentless: Demystifying LLM-based Software Engineering Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering