DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding

summary

Video file (mp4)

The gist

This paper introduces DeepDiscovery, a task-level repository-understanding method designed to address the limitations of large language models in navigating large industrial codebases.

In short

DeepDiscovery is a two-stage 'Location and Inference' framework designed to help AI understand entire code repositories instead of isolated snippets. The hosts discuss how its metadata-first approach manages token budgets and note that it achieved a 78.6% solve rate on the SWE-bench Verified benchmark without requiring offline preprocessing.

Key concepts

Location and Inference Framework
A two-stage process where the 'Location' stage finds starting points via high-confidence anchors like APIs, and the 'Inference' stage expands outward by tracing logic flows and implicit links, such as dependency injection, to map out the repository.
Metadata-first approach
A design choice used to manage token budgets by prioritizing a structural view of the code. The framework only retrieves a file's full text if its metadata is insufficient to explain its relevance to the specific task at hand.
SWE-bench Verified benchmark
A testing standard used to measure an AI's ability to solve real-world software issues. DeepDiscovery demonstrated significant performance on this benchmark, achieving a 78.6% solve rate, which was an 8.2 percentage point increase over the baseline.

Terminology used across episodes

This episode discusses

The paper

DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding · Read on arXiv

AMAP · Alibaba Group · Nanyang Technological University · Nanjing University · University of Cambridge

Large language models have shown strong performance on software engineering (SE) tasks, yet understanding large industrial repositories remains challenging. Existing methods often retrieve only local fragments and fail to recover the broader task-relevant context needed for complex repository-level tasks. We present DeepDiscovery, a task-level repository-understanding method for large industrial codebases. DeepDiscovery uses a two-stage Location--Inference framework to localize high-confidence task anchors and recover broader task-relevant context over multi-relational repository structure under budget constraints. Across controlled method-level evaluation, organization-internal industrial repository-understanding scenarios, and end-to-end evaluation on SWE-bench Verified, DeepDiscovery consistently improves task-relevant file recovery and downstream SE performance. On 27 medium-scale tasks, DeepDiscovery achieves the best file recovery quality among five representative baselines without offline preprocessing. On organization-internal industrial tasks from a production-scale integrated codebase ecosystem, including 27 medium-scale tasks and 40 large-scale tasks, DeepDiscovery improves Full Recall Rate across multiple AI coding systems, with absolute gains ranging from 1.6 to 9.2 percentage points on large subprojects and from 2.5 to 7.4 percentage points on medium-scale subprojects. In a controlled end-to-end evaluation on SWE-bench Verified, a system equipped with DeepDiscovery achieves a 78.6% Solve Rate, outperforming the corresponding baseline by 8.2 percentage points. These results suggest that stronger task-level repository understanding can improve coding-agent performance on complex SE tasks.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: The paper was written by Jiawei He, Weisong Sun, Mengyu Shi, Jie Jia, Tong Bian et al. from AMAP and Alibaba Group and Nanyang Technological University and Nanjing University and University of Cambridge.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are looking at a fascinating new paper called "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: This work comes from a massive collaboration between researchers at Alibaba, Nanyang Technological University, Nanjing University, and the University of Cambridge.

Tom: Jane, why is this specific title so important for the current state of AI?

Jane: Most AI models today struggle because they only look at small snippets of code instead of the whole project.

Tom: So they are basically looking through a keyhole?

Jane: That's a great way to put it, as they miss the connections between different files.

Tom: I imagine that makes fixing complex bugs nearly impossible.

Jane: It does, because you might find the error but miss the configuration file that actually causes it.

Tom: That sounds like a massive headache for anyone trying to use AI for real engineering.

Lu: I see it as an AI seeing a single gear without knowing how the entire machine functions.

Meng: I wonder if this framework can actually handle the scale of a production codebase with twenty-five thousand files.

Lalam: This approach helps us move toward a future where technology understands the purpose behind our code rather than just the text.

Tom: It sounds like they are trying to give the AI a sense of direction within a massive maze.

Jane: They really are, and it starts with how they identify where to begin looking.

Tom: Let's look at how they actually map out those paths.

Summary: Jane: To move beyond those small snippets, the researchers developed a two-stage process called Location and Inference.

Tom: Does the Location stage act as a way to find the right starting points?

Jane: Exactly, they look for these high-confidence anchors like specific APIs or important configuration files.

Tom: Once they have those anchors, how do they move through the rest of the repository?

Jane: That is where the Inference stage takes over to expand outward from those initial points.

Lu: I find it brilliant that they include implicit links like dependency injection instead of just standard imports.

Tom: So it's more like how a human developer traces a logic flow?

Lu: It really is, because humans know that a change in one module might trigger an event in another.

Meng: I was reading about their metadata-first approach to manage the token budget.

Jane: That's a very practical part of the design, Meng.

Meng: It makes sense because they only pull in the full text of a file if the metadata isn't enough to explain it.

Lalam: By prioritizing that structural view, they are teaching AI to build a mental map before reading every line.

Tom: That seems like it would prevent the model from getting overwhelmed by useless data.

Jane: It definitely helps keep the focus on what is actually relevant to the task at hand.

Tom: Let's see if this method actually delivers better results than what we are using now.

Improvements: Tom: The performance numbers in this paper are quite striking, especially when you look at the solve rates.

Jane: They reported a seventy-eight point six percent solve rate on the SWE-bench Verified benchmark.

Tom: That is an absolute gain of eight point two percentage points over their baseline!

Jane: It shows that providing better context directly translates to solving more real-world issues.

Lu: I was particularly impressed that they validated this on a massive ecosystem with over two million lines of code.

Meng: Did they compare their speed to the traditional RAG methods we see so often?

Lu: They did, and they noted that DeepDiscovery doesn't require any offline preprocessing at all.

Meng: That is a huge relief for engineers because we don't want to wait hours for an index to update every time we commit code.

Lalam: This reliability helps build a culture where AI can be trusted with much more complex software creation.

Tom: It sounds like they have found a way to make repository understanding both deep and fast.

Jane: They really have, especially since it works across different types of AI coding systems.

Tom: We should wrap this up by looking at the bigger picture for the industry.

Conclusion: Tom: We have spent our time today discussing "DeepDiscovery: A Location-Inference Framework for Task-Level Repository Understanding".

Jane: It feels like we are witnessing a shift from AI that just reads code to AI that actually understands software architecture.

Tom: Do you think this will change the way we build large-scale systems?

Jane: I believe so, because the gap between simple chatbots and true coding agents is finally closing.

Lu: I can imagine this leading to autonomous architects that can simulate how a whole system will react to a change.

Meng: As long as it stays efficient and respects our latency requirements, I am all for it.

Lalam: When our tools finally grasp the underlying structure of our work, it allows humans to focus on higher-level creativity.

Tom: That is a perfect note to end on, Lalam.

Jane: Thanks to the whole team for joining us to break down this incredible research.

Tom: We will see you all next time!

More episodes

← Home