Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale".
Tom: Language models can act as corpus-scale retrievers, but their performance degrades significantly when scaling to million-token corpora and generalizing beyond training sizes.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, moving on, let’s get into a quick rundown of what the paper "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale" is actually about. Essentially, the main thesis is that language models can act as corpus-scale retrievers by conditioning on an in-context corpus and directly generating an answer, which is an intriguing alternative to older vector-based retrieval methods.
Jane: That’s a really simple way to put it—instead of searching a database for documents first, you just give the model the context and have it generate what you need. But they are focusing on pushing that capability to its limits with massive contexts.
Lu: The paper claims that prior work has mostly focused on smaller systems or reranking tasks, leaving corpus-scale in-context retrieval largely unexplored when dealing with million-token corpora and length generalization far beyond training sizes.
Meng: So they are specifically looking at the two hardest things to achieve: scaling up to corpora containing millions of tokens, and generalizing the model's ability to handle contexts much longer than it was ever seen during its initial training.
Lalam: It matters because if we can solve these scaling problems, it opens up a whole new way for AI systems to access and utilize knowledge that is currently too vast or too long for standard methods.
Tom: They introduce BLOCKSEARCH, which is a 0 point 6B parameter LM retriever designed specifically to handle these exact challenges, and they show it can generalize up to ten times beyond its training regime through architectural changes <ref:2607.01538#pg0>.
Jane: That generalization claim is pretty big; it means the model isn't just good at what it saw in training but has learned a transferable skill for handling much larger inputs.
Lu: To achieve this, they build upon prior ICR architectures by adding randomized per-document identifier codes, utilizing in-batch negative training, and incorporating an on-policy auxiliary loss that trains on the model’s own rolled-out digit prefixes.
Meng: Those are specific training adjustments aimed at making the model more robust when it sees context it hasn't seen before, which is a very practical approach to improving reliability.
Lalam: I think those methods show a concrete path to making these large models more versatile tools rather than just narrow systems that work perfectly only in their immediate training environment.
Tom: The core of their investigation is tracing the failure modes of in-context retrieval as the corpus size grows, specifically identifying "attention dilution" as the primary bottleneck that causes recall to approach zero near million-token contexts.
Jane: So, they aren't just saying it gets worse; they are showing a detailed mechanism for *why* it gets worse—that irrelevant documents start overwhelming the scores in the attention mechanism.
Lu: They decompose this collapse into two measurable components: a per-head, per-layer measure of attention recall and the generation recall under those extreme conditions.
Meng: Measuring that specific failure mode is key; it moves the discussion from just observing poor results to understanding the underlying mechanics of why those results are poor.
Lalam: Understanding this dilution effect gives us a much clearer picture of where we need to focus our development efforts for more robust AI systems when they encounter massive amounts of text.
Conclusion: Tom: So, wrapping up this discussion on "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale," the authors are essentially proving that LMs can perform meaningful retrieval far beyond their normal training regimes, constrained mainly by this attention dilution effect.
Jane: It’s a significant finding because it moves the focus from just building bigger models to engineering smarter ways to make the existing architecture handle massive data and long contexts effectively.
Lu: The implication here is that the path forward involves developing mechanisms, like length-aware softmax modifications, which actively adjust how attention mass is distributed across the context tokens.
Meng: From a practical standpoint, this means that future retrieval systems won't just rely on raw model size; they will need to integrate these specific architectural tweaks to ensure reliability when dealing with massive real-world documents.
Lalam: For culture and application, this suggests we can build AI tools that are far more capable of handling vast knowledge bases, which could lead to new kinds of complex problem-solving capabilities we haven't even conceived yet.
Tom: The paper really highlights that the limitations aren't just about the context window length anymore; they are about the fundamental way attention behaves when scaled up.
Jane: Exactly; it shows that understanding these failure modes, like softmax dilution, is a necessary step before we can build truly scalable retrieval systems for massive datasets.
Lu: They’ve pointed toward solutions that involve applying techniques like attention sinks to mitigate this dilution effect under large N scenarios, showing that these mechanisms might hold the key to overcoming the collapse in performance.
Meng: It tells me we need engineers to start thinking about these kinds of dynamic adjustments as standard components when designing new retrieval pipelines, rather than just adding them as experimental fixes later.
Lalam: This research gives us a lot of confidence that AI systems have the potential to operate in ways that are far more powerful and flexible than our current limited training contexts allow.
UC Berkeley · UT Austin
cs.CL
Submitted: 2026-07-01
Updated: 2026-10-03
Importance score: 90/100
The gist: Language models can act as corpus-scale retrievers, but their performance degrades significantly when scaling to million-token corpora and generalizing beyond training sizes.
Key concepts
- Attention Dilution
- This is the main problem encountered when using LLMs for retrieval over very large datasets. As the number of documents grows, the model's attention becomes diluted, meaning its ability to focus on a single, relevant gold document weakens significantly. This causes retrieval recall to drop sharply near million-token scales.
- BLOCKSEARCH
- BLOCKSEARCH is a specific retriever designed for handling huge corpora and long documents. It uses two key strategies: inserting unique document codes into the text and using block-sparse attention, where tokens only attend within their own document blocks while the query attends across the whole corpus.
- Length-aware Softmax Modifications
- These are techniques used to fix attention dilution by changing how scores are calculated during retrieval. Methods like Multiplicative Score Rescaling (SSMax) multiply pre-softmax scores by a factor related to the corpus size, effectively widening the gap between the score of a correct document and incorrect distractors.
Terminology
Summary
Language models can act as corpus-scale retrievers, but their performance degrades significantly when scaling to million-token corpora and generalizing beyond training sizes. This work presents the first systematic study of in-context retrieval on these scales, introducing modifications that mitigate attention dilution
to show that LMs can perform meaningful retrieval far beyond their nominal training regimes.
BLOCKSEARCH: Model and Million-Token Evaluation
The researchers introduce BLOCKSEARCH, a 0.6B-parameter LM retriever designed to handle the challenges of million-token corpora (C1) and length generalization (C2). To address the risk of overfitting to absolute positions, they insert each document as Doc [code]: [text]
with a randomly drawn code per training step. Furthermore, they utilize block-sparse attention,
where document tokens attend causally only within their own block, and the query block attends over the full corpus. The training involves using the ReLabeled Hard Negatives (RLHN) version of BEIR and an on-policy auxiliary loss to mitigate exposure bias during decoding.
Understanding Large-N Deterioration
The study identifies attention dilution
as the primary bottleneck at million-token scale, which causes retrieval recall to approach zero near million-token contexts. They decompose this collapse into two parts: a perhead, per-layer measure of attention recall
and the generation recall.
Specifically, they show that while the pre-softmax score is preserved,
the aggregate contribution of irrelevant documents to the softmax denominator grows faster with corpus size, causing the normalized attention mass on the gold document to collapse.
This is evidenced by Table 9, which shows that for L19 (the first-digit emission position), while gold's pre-softmax score drops by ∼ 3 logit units,
the noise gap widens by ∼ 3.5,
causing the per-head gold mass to collapse by ∼ 150×.
Addressing Attention Dilution Problem
To mitigate attention dilution, two main techniques are explored:
-
Length-aware softmax modifications: These include an additive sink, which appends a learned constant to the softmax denominator, and Multiplicative Score Rescaling (SSMax), which multiplies pre-softmax scores by
s · log N
to widen the gap between gold and distractor scores. The additive sink is shown to be equivalent to multiplying the standard softmax by a sigmoid gate dependent on layer logits. -
Document-level sparse attention: This involves performing a
doc-level routing step at L16,
where only tokens of the top-B shortlisted documents participate in attention, which they termBLOCKSEARCHrouting.
Out-of-Distribution Generalization and Final Results
The proposed modifications are evaluated on benchmarks requiring different similarity notions, such as LIMIT, where dense retrieval struggles. The combination of techniques—specifically BLOCKSEARCH-SSMax-routing—demonstrates superior performance. On LIMIT, this method significantly outperforms the Qwen3-dense baseline by nearly 3×. While the attention ceiling R any19 stays at 1.00 across the entire sweep,
indicating that at least one head still ranks a gold document first,
the final code-generation readout
recovers this signal more effectively than simpler methods like BLOCKSEARCH or BLOCKSEARCH-sink, suggesting that attention dilution is a fundamental challenge for scalable in-context retrieval. The paper concludes that LMs can perform meaningful retrieval far beyond their nominal training regime, constrained primarily by attention dilution.
Key Findings Summary
(The paper enumerates the following key findings and results):
-
BLOCKSEARCH matches dense retrievers at smaller corpus sizes (e.g., 95.8% vs. 95.5% on MS MARCO at 45k tokens).
-
BLOCKSEARCH generalizes up to
10× beyond its training context.
-
The collapse of recall at N=10,000 is traced to the
attention dilution effect.
-
Length-aware modifications like SSMax close the gap to Qwen3-dense on benchmarks like MS MARCO (e.g., 33.8% vs 20.2% at N=5k).
-
On out-of-distribution tasks like LIMIT, BLOCKSEARCH significantly outperforms dense retrieval by nearly 3× (e.g., R@1 of 0.149 vs Qwen3-dense's 0.176 at N=5,000).
-
The
attention ceiling R any19
remains at 1.0 across all tested corpus sizes and methods in the LIMIT evaluation, confirming that the retrieval signal persists internally even when the final readout fails. -
SSMax+routing achieves strong results on MS MARCO (20.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on this research, and what those improved systems could achieve:
) Improved System Capabilities: Corpus-Scale In-Context Retrieval (ICR) for LLMs
The core improvement lies in moving beyond traditional vector search by enabling Large Language Models (LLMs) to act as direct, corpus-scale retrievers without relying on a separate retrieval step.
-
// Architecture: Implement BLOCKSEARCH or its variants (BLOCKSEARCH-sink, BLOCKSEARCH-SSMax) for LLM retrieval tasks.
-
// Training: Fine-tune the LLM (e.g., Qwen/Qwen3) using the specialized training regimen described, including in-batch contrastive learning with hard negatives and an on-policy auxiliary loss to mitigate exposure bias during code generation.
-
// Length Generalization: The improved system can retrieve relevant information from context sizes up to 10 times larger than its nominal training context length (i.e., generalizing beyond the 32k token limit of the backbone).
) Specific System Improvements and Performance Gains:
-
// Enhanced Retrieval on Massive Corpora:
-
The system can reliably retrieve relevant documents from corpora containing millions of tokens (up to 1M tokens in testing), matching or exceeding dense retrieval performance on widely studied benchmarks like MS MARCO and NQ at these scales, which is currently impossible for standard models.
-
// Out-of-Distribution (OOD) Robustness:
-
The system exhibits superior generalization to completely new similarity notions, such as lexical similarity (e.g., the LIMIT benchmark), where dense retrieval methods fail, achieving up to 3× higher performance compared to dense baselines.
-
// Mitigating Attention Dilution Bottleneck:
-
The improved system overcomes the primary scaling bottleneck identified—attention dilution—by employing length-aware modifications (like additive sinks or multiplicative score rescaling) or document-level sparse attention routing. This ensures that the retrieval signal remains strongly present in the attention mechanism even as context grows to millions of tokens.
-
// Superior Performance vs. Concurrent Models:
-
The system can outperform concurrent, much larger models trained for long contexts (like MSA-4B) while being significantly smaller (e.g., 7× fewer parameters), demonstrating better efficiency in achieving corpus-scale retrieval capabilities compared to existing state-of-the-art methods.
-
// Versatile Retrieval Modalities:
-
The system is effective across different similarity paradigms, including lexical tasks (LIMIT) and oblique reasoning tasks (OBLIQ), showing the potential for ICR to be a more powerful general retrieval mechanism than embedding-based search alone.
) Summary of What the Improved AI System Can Do:
The resulting improved AI system would function as an Intelligent Corpus Navigator
capable of:
-
// Large-Scale Knowledge Extraction: Instantly finding specific, highly relevant passages within vast, unstructured documents (millions of tokens) based on complex queries.
-
// Robust Long-Context Reasoning: Maintaining high retrieval accuracy even when the input context is significantly larger than the model was trained to handle, ensuring reliability in real-world, large-scale data analysis or document synthesis tasks.
-
// Cross-Domain Retrieval Excellence: Successfully performing retrieval tasks that require semantic understanding beyond simple keyword matching (e.g., identifying passages with analogous reasoning patterns), outperforming specialized embedding models in these niche areas.
Sources
- Can Long-Context Language Models Subsume Retrieval, RAG, SQL, and More?
- On the Theoretical Limitations of Embedding-Based Retrieval
- Scalable In-context Ranking with Generative Models
- Scalable-Softmax Is Superior for Attention
- Understanding and Improving Length Generalization in Hierarchical Sparse Attention Models
- Softmax is not Enough (for Sharp Size Generalisation)
- LUCID: Attention with Preconditioned Representations
- MoBA: Mixture of Block Attention for Long-Context LLMs
- Optimizing Mixture of Block Attention
- MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens
- Block Sparse Flash Attention
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- Efficient Length-Generalizable Attention via Causal Retrieval for Long-Context Language Modeling
- Efficient Streaming Language Models with Attention Sinks
- gpt-oss-120b & gpt-oss-20b Model Card
- Drowning in Documents: Consequences of Scaling Reranker Inference
- Block-Attention for Efficient Prefilling
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering