SHERLOC: Structured Diagnostic Localization for Code Repair Agents
cs.CL
Submitted: 2026-06-23
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 Main Conference. Code: https://github.com/NVIDIA-NeMo/Skills/tree/main/recipes/sherloc
Code: https://github.com/NVIDIA-NeMo/Skills
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing.
Terminology
Abstract
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have emerged, yet are still evaluated as file retrieval rather than actionable diagnosis, producing locations without the diagnostic context a repair agent needs. We introduce SHERLOC (Structured Hypothesis-driven Exploration and Reasoning for Localization), a training-free framework pairing a reasoning LLM with compact repository tools and self-recovery, without fine-tuning or multi-agent orchestration. SHERLOC reaches state-of-the-art localization across model scales: 84.33% accuracy@1 on SWE-Bench Lite and 81.27% recall@1 on SWE-Bench Verified; at 30B parameters, it matches or outperforms other agentic methods. Injecting our locations and diagnostic findings into repair agents yields an average +5.95 pp resolve-rate gain from the best SHERLOC result per setting on SWE-Bench Verified. SHERLOC cuts localization and total tokens by 36.7% and 23.1% on average.
Sources
- DeepSeek-V3 Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Issue Localization via LLM-Driven Iterative Code Graph Searching
- SWE-Debate: Competitive Multi-Agent Debate for Software Issue Resolution
- The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
- Tool-integrated Reinforcement Learning for Repo Deep Search
- Toolformer: Language Models Can Teach Themselves to Use Tools
- Qwen3 Technical Report
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
- ReAct: Synergizing Reasoning and Acting in Language Models
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Agentless: Demystifying LLM-based Software Engineering Agents
- OrcaLoca: An LLM Agent Framework for Software Issue Localization
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering