Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
cs.SE, cs.AI, cs.CL, cs.IR
Submitted: 2025-12-05
Updated: 2025-12-05
Comments: Accepted at LLM4Code Workshop, ICSE 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Bug localization in multi-repository microservice architectures is challenging due to the semantic gap between natural language bug reports and code, LLM context limitations, and the need to first
Terminology
Abstract
Bug localization in multi-repository microservice architectures is challenging due to the semantic gap between natural language bug reports and code, LLM context limitations, and the need to first identify the correct repository. We propose reframing this as a natural language reasoning task by transforming codebases into hierarchical NL summaries and performing NL-to-NL search instead of cross-modal retrieval. Our approach builds context-aware summaries at file, directory, and repository levels, then uses a two-phase search: first routing bug reports to relevant repositories, then performing top-down localization within those repositories. Evaluated on DNext, an industrial system with 46 repositories and 1.1M lines of code, our method achieves Pass@10 of 0.82 and MRR of 0.50, significantly outperforming retrieval baselines and agentic RAG systems like GitHub Copilot and Cursor. This work demonstrates that engineered natural language representations can be more effective than raw source code for scalable bug localization, providing an interpretable repository -> directory -> file search path, which is vital for building trust in enterprise AI tools by providing essential transparency.
Sources
- Towards Explorative IRBL: Combining Semantic Retrieval with LLM-driven Iterative Code Exploration
- Bridging Bug Localization and Issue Fixing: A Hierarchical Localization Framework Leveraging Large Language Models
- GraphCodeBERT: Pre-training Code Representations with Data Flow
- Neural Code Search Revisited: Enhancing Code Snippet Retrieval through Natural Language Intent
- LLM Agents Improve Semantic Code Search
- Issue Localization via LLM-Driven Iterative Code Graph Searching
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Rewriting the Code: A Simple Method for Large Language Model Augmented Code Search
- GraphCoder: Enhancing Repository-Level Code Completion via Code Context Graph-based Retrieval and Language Model
- Improving Source Code Similarity Detection Through GraphCodeBERT and Integration of Additional Features
- Code-Craft: Hierarchical Graph-Based Code Summarization for Enhanced Context Retrieval
- Source Code Summarization in the Era of Large Language Models
- Commenting Higher-level Code Unit: Full Code, Reduced Code, or Hierarchical Code Summarization
- Meta-RAG on Large Codebases Using Code Summarization
- Improving LLM-Based Fault Localization with External Memory and Project Context
- OrcaLoca: An LLM Agent Framework for Software Issue Localization
Related papers
- Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
- GitSkills: A Dataset of Agent Skills on GitHub
- SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces
- PackMonitor: Enabling Zero Package Hallucinations Through Decoding-Time Monitoring
- IntentCoding: Amplifying User Intent in Code Generation
- Incentives and Outcomes in Bug Bounties