OrcaLoca: An LLM Agent Framework for Software Issue Localization
Zhongming Yu, Hejia Zhang, Yujie Zhao, Hanxian Huang, Matrix Yao, Ke Ding, Jishen Zhao
University of California, San Diego · Intel Corporation
cs.SE, cs.AI
Submitted: 2025-10-10
Updated: 2026-08-20
Code: https://github.com/fishmingyu/OrcaLoca
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 77/100
The gist: OrcaLoca: An LLM Agent Framework for Software Issue Localization The paper introduces OrcaLoca, an LLM agent framework designed to address the significant challenge of software issue localization
Terminology
Summary
OrcaLoca: An LLM Agent Framework for Software Issue Localization
The paper introduces OrcaLoca, an LLM agent framework designed to address the significant challenge of software issue localization within Autonomous Software Engineering (ASE). While Large Language Models (LLMs) have advanced rapidly, localization—the ability to precisely identify and navigate to relevant code for resolving software engineering problems
—remains a crucial yet underexplored challenge.
Problem Statement and Challenges
The inherent complexity of large software repositories, such as the average SWE-bench codebase consisting of 3,010 files with around 438K lines of code, coupled with imprecise natural language requirements for user issues, makes effective localization difficult. The paper identifies three key challenges in existing LLM agent-based localization approaches:
-
Action Planning Inefficiencies: Previous methods relied
solely on LLMs for guidance,
resulting inunstable and redundant search behaviors.
-
Context Management Trade-offs: Existing approaches struggle to achieve both conciseness (reducing noise) and search space completeness (ensuring no critical details are omitted), often optimizing one at the expense of the other.
3** Context Accumulation:** Large repositories introduce noise, and existing frameworks merely concatenate all search results into the context,
which is insufficient to manage complexity.
Proposed Methodology: The OrcaLoca Framework
To overcome these limitations, OrcaLoca proposes an agent system built upon three key components:
-
Priority-Based Scheduling for LLM-Guided Actions: This system incorporates
priority queues and LLM-guided action generation
to dynamically reorder actions based on their contextual relevance and urgency, solving the shortcomings of previous systems thatlacked effective action management.
-
Action Decomposition with Relevance Scoring: This method resolves the trade-off between conciseness and completeness by decomposing high-level actions (such as class or file skeletons) into
finer-grained sub-actions.
These sub-actions are then evaluated and ranked using a multi-agent workflow, ensuringcomprehensive exploration while avoiding noise and redundancy.
-
Distance-Aware Searched Context Pruning: The framework utilizes a context manager that dynamically prunes the searched context. This pruning algorithm leverages
a node distance heuristic within the graph-oriented codebase
to filter out irrelevant data, ensuring that exploration remains focused and aligned with the bug localization.
Technical Implementation
OrcaLoca constructs a CodeGraph, a graph-based representation G = (V,E) of the codebase, which facilitates indexing and searching. This CodeGraph integrates two primary edge types: e 1 (containment/hierarchy) and e 2 (reference/function calls).
The agent follows a reason-and-act workflow
where the state transition is managed by the LLM, and actions are managed by an Action Scheduler Queue (ASQ). The system employs several mechanisms to enhance robustness:
-
Redundancy Elimination: A mechanism is implemented to ensure that
redundant actions are avoided, enhancing efficiency and preventing unnecessary exploration.
-
Disambiguation: To handle ambiguities common in large repositories (such as function overrides or inherited classes), a robust disambiguation mechanism is implemented using an inverted index.
-
Context Pruning (CM): The Context Manager evaluates each search result (v SR) based on its proximity to the potential bug locations (P B). The pruning process calculates the
average shortest path distance
between v SR and candidate nodes in P B, retaining only those SR entries linked to valid search queries.
Evaluation and Results
The performance of OrcaLoca was evaluated on the SWE-bench Lite dataset against 17 different approaches. The key metrics used were Resolved Rate, Function Match Rate, File Match Rate, and Function Match Precision.
OrcaLoca achieved a new open-source State-Of-The-Art (SOTA) in the function match rate with 65.33% (190 out of 300) and a file match rate of 83.33% (250 out of 300).
Furthermore, by integrating the patch generation component from Agentless-1.5, OrcaLoca demonstrated an impact on the final resolved rate, achieving a 6.67 percentage points improvement in function match rate and a 6.33 percentage points gain in Resolved Rate.
The study also conducted an ablation analysis on SWE-bench Common, confirming that removing any of the core methods (priority scheduling, decomposition, or context pruning) resulted in a performance drop of approximately 3–5 percentage points.
Improvements for AI systems
The following improvements are derived from the principles of OrcaLoca. These modifications move beyond simple sequential Chain-of-Thought (CoT) prompting to create a robust, heuristic-driven agentic workflow suitable for industrial application.
We propose integrating three core modules—Priority Scheduling, Action Decomposition, and Context Pruning—into the agent's operational framework. These components address the fundamental limitations of current LLM-based agents (redundancy, scope ambiguity, and context overload).
Instead of relying solely on the LLM to generate a sequence of actions, the system must implement an external Action Search Queue (ASQ) governed by priority heuristics.
-
Mechanism: Maintain a dynamic priority queue based on contextual relevance and action urgency (C ak). Actions are not executed in LLM-suggested order; they are prioritized based on their known connections to the potential bug locations (PB) and their frequency of appearance in the problem statement.
-
Anti-Redundancy Logic: The system tracks previously executed actions and increments a counter (C ak). If an action is repeated, its priority increases, ensuring that the LLM focuses on verifying or refining critical areas rather than repeating exploration.
The agent must transition from monolithic queries to a recursive, granular search strategy.
-
Mechanism: When faced with a large entity (e.g, a massive class or file), the system automatically triggers an Action Decomposer. This sub-agent breaks down the large entity into smaller, functional units (methods/functions).
-
Relevance Scoring: Each decomposed sub-action is then subjected to a multi-agent relevance scoring mechanism (Top-k selection). Only the most relevant sub-actions are pushed back into the ASQ, ensuring comprehensive coverage while discarding noise.
The system must move away from simply concatenating all search results into the LLM's context window.
-
Mechanism: The agent requires a pre-processed CodeGraph (G), which maps all code entities to graph nodes, encoding both containment (e 1) and reference (e 2) relationships.
-
Pruning Logic: A Context Manager (CM) is implemented using a distance-aware heuristic. For every new search result (SR), the CM calculates its shortest path distance to the potential bug locations (PB) within the CodeGraph: d(SR, PB) = (d(v SR, v), d(v, v SR)). Only SR entries with a high relevance score (low distance) are retained in the LLM context.
The integration of these methods transforms the AI system from a search-and-guess
mechanism into an efficient, structural, and highly targeted problem solver.
-
Structural Autonomy: The system can autonomously navigate massive code repositories (e.g., those exceeding 3,000 files) without being overwhelmed by noise or losing track of the core issue.
-
High Efficiency and Low Latency: By utilizing a dynamic priority queue and context pruning, the system ensures that the most critical and relevant code sections are analyzed first, drastically reducing redundant searches (demonstrated by significant token cost reduction).
-
Guaranteed Granularity: The decomposition module ensures that even complex classes or files are thoroughly investigated at a granular level, preventing the loss of subtle bugs hidden within large code blocks.
-
Optimized Resolution Path: The combination of structured planning and relevance-based context filtering leads directly to higher accuracy in identifying the root cause, translating to a measurable increase in the final Resolved Rate compared to state-of-the-art models.
Sources
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- Evaluating Large Language Models Trained on Code
- A Quantitative and Qualitative Evaluation of LLM-Based Explainable Fault Localization
- ContrastRepair: Enhancing Conversation-Based Automated Program Repair via Contrastive Test Case Pairs
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- MarsCode Agent: AI-native Automated Bug Fixing
- Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration
- Large Language Models in Fault Localisation
- Agentless: Demystifying LLM-based Software Engineering Agents
- RepoGraph: Enhancing AI Software Engineering with Repository-level Code Graph
- Qwen3 Technical Report
- HyperAgent: Generalist Software Engineering Agents to Solve Coding Tasks at Scale
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- AgentFL: Scaling LLM-based Fault Localization to Project-Level Context
- ReAct: Synergizing Reasoning and Acting in Language Models
- LLaMA: Open and Efficient Foundation Language Models
- Executable Code Actions Elicit Better LLM Agents
- OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties