DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding

arXiv:2606.22906 · cs.SE, cs.AI · Submitted 2026-06-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: The paper was written by Jiawei He, Weisong Sun, Mengyu Shi, Jie Jia, Tong Bian et al. from AMAP and Alibaba Group and Nanyang Technological University and Nanjing University and University of Cambridge.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are looking at a fascinating new paper called "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: This work comes from a massive collaboration between researchers at Alibaba, Nanyang Technological University, Nanjing University, and the University of Cambridge.

Tom: Jane, why is this specific title so important for the current state of AI?

Jane: Most AI models today struggle because they only look at small snippets of code instead of the whole project.

Tom: So they are basically looking through a keyhole?

Jane: That's a great way to put it, as they miss the connections between different files.

Tom: I imagine that makes fixing complex bugs nearly impossible.

Jane: It does, because you might find the error but miss the configuration file that actually causes it.

Tom: That sounds like a massive headache for anyone trying to use AI for real engineering.

Lu: I see it as an AI seeing a single gear without knowing how the entire machine functions.

Meng: I wonder if this framework can actually handle the scale of a production codebase with twenty-five thousand files.

Lalam: This approach helps us move toward a future where technology understands the purpose behind our code rather than just the text.

Tom: It sounds like they are trying to give the AI a sense of direction within a massive maze.

Jane: They really are, and it starts with how they identify where to begin looking.

Tom: Let's look at how they actually map out those paths.

Summary: Jane: To move beyond those small snippets, the researchers developed a two-stage process called Location and Inference.

Tom: Does the Location stage act as a way to find the right starting points?

Jane: Exactly, they look for these high-confidence anchors like specific APIs or important configuration files.

Tom: Once they have those anchors, how do they move through the rest of the repository?

Jane: That is where the Inference stage takes over to expand outward from those initial points.

Lu: I find it brilliant that they include implicit links like dependency injection instead of just standard imports.

Tom: So it's more like how a human developer traces a logic flow?

Lu: It really is, because humans know that a change in one module might trigger an event in another.

Meng: I was reading about their metadata-first approach to manage the token budget.

Jane: That's a very practical part of the design, Meng.

Meng: It makes sense because they only pull in the full text of a file if the metadata isn't enough to explain it.

Lalam: By prioritizing that structural view, they are teaching AI to build a mental map before reading every line.

Tom: That seems like it would prevent the model from getting overwhelmed by useless data.

Jane: It definitely helps keep the focus on what is actually relevant to the task at hand.

Tom: Let's see if this method actually delivers better results than what we are using now.

Improvements: Tom: The performance numbers in this paper are quite striking, especially when you look at the solve rates.

Jane: They reported a seventy-eight point six percent solve rate on the SWE-bench Verified benchmark.

Tom: That is an absolute gain of eight point two percentage points over their baseline!

Jane: It shows that providing better context directly translates to solving more real-world issues.

Lu: I was particularly impressed that they validated this on a massive ecosystem with over two million lines of code.

Meng: Did they compare their speed to the traditional RAG methods we see so often?

Lu: They did, and they noted that DeepDiscovery doesn't require any offline preprocessing at all.

Meng: That is a huge relief for engineers because we don't want to wait hours for an index to update every time we commit code.

Lalam: This reliability helps build a culture where AI can be trusted with much more complex software creation.

Tom: It sounds like they have found a way to make repository understanding both deep and fast.

Jane: They really have, especially since it works across different types of AI coding systems.

Tom: We should wrap this up by looking at the bigger picture for the industry.

Conclusion: Tom: We have spent our time today discussing "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: It feels like we are witnessing a shift from AI that just reads code to AI that actually understands software architecture.

Tom: Do you think this will change the way we build large-scale systems?

Jane: I believe so, because the gap between simple chatbots and true coding agents is finally closing.

Lu: I can imagine this leading to autonomous architects that can simulate how a whole system will react to a change.

Meng: As long as it stays efficient and respects our latency requirements, I am all for it.

Lalam: When our tools finally grasp the underlying structure of our work, it allows humans to focus on higher-level creativity.

Tom: That is a perfect note to end on, Lalam.

Jane: Thanks to the whole team for joining us to break down this incredible research.

Tom: We will see you all next time!

AMAP · Alibaba Group · Nanyang Technological University · Nanjing University · University of Cambridge

cs.SE, cs.AI

Submitted: 2026-06-22

Updated: 2026-09-14

Comments: 12 pages, 3figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: This paper introduces DeepDiscovery, a task-level repository-understanding method designed to address the limitations of large language models in navigating large industrial codebases.

Key concepts

Location and Inference Framework
A two-stage process where the 'Location' stage finds starting points via high-confidence anchors like APIs, and the 'Inference' stage expands outward by tracing logic flows and implicit links, such as dependency injection, to map out the repository.
Metadata-first approach
A design choice used to manage token budgets by prioritizing a structural view of the code. The framework only retrieves a file's full text if its metadata is insufficient to explain its relevance to the specific task at hand.
SWE-bench Verified benchmark
A testing standard used to measure an AI's ability to solve real-world software issues. DeepDiscovery demonstrated significant performance on this benchmark, achieving a 78.6% solve rate, which was an 8.2 percentage point increase over the baseline.

Terminology

Summary

This paper introduces DeepDiscovery, a task-level repository-understanding method designed to address the limitations of large language models in navigating large industrial codebases. By shifting the focus from simple fragment retrieval to task-relevant context recovery, the framework enables more effective reasoning over complex software engineering tasks while operating under realistic freshness, budget, and deployment constraints.

The core challenge

Large industrial repositories present significant difficulties due to massive numbers of files, complex dependencies, mixed artifact types, and continuously evolving organization. Existing methods often retrieve only local fragments and fail to capture the broader context required for complex tasks. Semantic retrieval may miss entities that are structurally important yet lexically less salient, while static dependency expansion struggles with implicit links such as configuration registration or dependency injection. Furthermore, many existing approaches rely on prebuilt vector indexes or static graphs that are costly to maintain in environments where repository freshness is a critical concern.

How it works

DeepDiscovery utilizes a two-stage Location–Inference framework to identify high-confidence task anchors and expand them into a broader implementation path. The Location stage narrows the search space through three steps:

  • Environment-Aware Analysis: Inferring the technology stack, module scope, and artifact-role priors.

  • Adaptive Repository Compression: Adjusting the granularity of the repository view based on repository scale, structural complexity, and context budget.

  • Rule-guided Anchor Localization: Ranking candidate entities using semantic similarity, structural summaries, rule-template matches, and task-conditioned priors.

The Inference stage then expands from these anchors over a multi-relational repository graph. This graph integrates three distinct sources of structure:

  1. Explicit dependencies, such as imports, calls, and inheritance.

  2. Implicit task-relevant links, such as configuration-to-code mappings and registration sites.

  3. Organizational relations, including folder containment and module boundaries.

Context construction and efficiency

To manage computational resources, the system employs budget-aware decision-making. During expansion, it prioritizes actions using a Priority score that balances the Gain in task relevance against the Cost of exploration and token usage. Once entities are selected, DeepDiscovery utilizes a metadata-first context construction approach. This method aims to preserve structural coverage while loading full text only when necessary, using a task-time promotion score to decide whether to provide a lightweight metadata view or the full-text content.

Experimental results

Evaluations across method-level, system-level, and end-to-end benchmarks demonstrate that DeepDiscovery consistently improves task-relevant file recovery. Key findings include:

  • On 27 medium-scale tasks, it achieves the best file recovery quality among five representative baselines without offline preprocessing.

  • In industrial settings, it improves the Full Recall Rate across multiple AI coding systems, with absolute gains ranging from 1.6 to 9.2 percentage points on large subprojects.

  • On the SWE-bench Verified benchmark, a system equipped with DeepDiscovery achieves a 78.6% Solve Rate, outperforming the corresponding baseline by 8.2 percentage points.

Improvements for AI systems

1. Transition from Fragment-Based RAG to Multi-Relational Implementation-Path Recovery

  • The Improvement: Replace standard semantic vector retrieval with a multi-relational expansion engine. This engine must integrate three distinct relation types: explicit dependencies (calls, imports, inheritance), implicit task-relevant links (Dependency Injection wiring, configuration-to-code mappings, event/callback bindings, and test-to-implementation bridges), and organizational relations (folder hierarchies, module boundaries, and physical proximity).

  • System Capability: The AI can navigate non-textual logic chains. Instead of just finding files that look like the query, the system can follow a configuration setting to the service it registers, then to the interface it implements, and finally to the test that validates it, even if those files share no lexical similarity.

2. Implementation of a Two-Stage Location–Inference Workflow

  • The Improvement: Implement a staged discovery process. Stage 1 (Location) performs environment-aware analysis (inferring tech stacks and module scope) and rule-guided anchor localization. Stage 2 (Inference) performs budgeted expansion from those high-confidence anchors.

  • System Capability: The system can perform surgical repository exploration. It avoids the high latency and token cost of global searches by first narrowing the search space to high-confidence entry points (anchors) before attempting to map out the broader implementation path.

3. Adaptive Repository Compression & Metadata-First Context Construction

  • The Improvement: Replace the flat list of files context injection with a hierarchical, metadata-driven approach. The system should construct Metadata Cards (summaries, structural roles, and directory context) as the primary view, using a Metadata-Promotion Threshold to trigger the loading of Full-Text Cards only when an entity is statistically likely to require deep inspection for reasoning.

  • System Capability: The AI can reason over massive (multi-million LOC) repositories within a standard context window. It maintains a high-level map of the project structure while selectively zooming in on the specific lines of code necessary for a task, maximizing information density and reducing lost-in-the-middle reasoning errors.

4. Rule-Guided Hybrid Anchor Localization

  • The Improvement: Augment embedding-based similarity scoring with a multi-signal scoring function: Score = alpha(Semantic) + beta(Structural Alignment) + gamma(Rule Template Matches) + delta(Artifact Role Priors). This includes injecting naming patterns, directory conventions, and framework-specific keywords (e.g., @Service, config.yaml, test*.py) into the localization logic.

  • System Capability: The system can identify critical entry points that semantic search misses, such as a specific configuration file, a registration site, or a business logic interface, by recognizing the idioms and roles of the files within the specific technology stack.

5. Online, Freshness-Oriented Context Discovery

  • The Improvement: Shift away from heavy, offline-preprocessed artifacts (static vector indexes and pre-built knowledge graphs) toward a lightweight, online discovery mechanism that utilizes adaptive compression and rule-based extraction at task-time.

  • System Capability: The AI agent becomes effective in high-velocity industrial environments. It can immediately work on new branches, recent commits, or evolving configurations without waiting for a time-consuming re-indexing or graph-reconstruction phase, ensuring the context is always synchronized with the current repository state.

Abstract

Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often retrieve only local fragments and fail to recover the broader task-relevant context needed for complex repository-level tasks. We present DeepDiscovery, a task-level repository-understanding method for large industrial codebases. DeepDiscovery uses a two-stage Location--Inference framework to localize high-confidence task anchors and recover broader task-relevant context over multi-relational repository structure under budget constraints. Across controlled method-level evaluation, organization-internal industrial repository-understanding scenarios, and end-to-end evaluation on SWE-bench Verified, DeepDiscovery consistently improves task-relevant file recovery and downstream SE performance. On 27 medium-scale tasks, DeepDiscovery achieves the best file recovery quality among five representative baselines without offline preprocessing. On organization-internal industrial tasks from a production-scale integrated codebase ecosystem, including 27 medium-scale tasks and 40 large-scale tasks, DeepDiscovery improves Full Recall Rate across multiple AI coding systems, with absolute gains ranging from 1.6 to 9.2 percentage points on large subprojects and from 2.5 to 7.4 percentage points on medium-scale subprojects. In a controlled end-to-end evaluation on SWE-bench Verified, a system equipped with DeepDiscovery achieves a 78.6% Solve Rate, outperforming the corresponding baseline by 8.2 percentage points. These results suggest that stronger task-level repository understanding can improve coding-agent performance on complex SE tasks.

Sources

Related papers