Information Abundance Paradox: Long-Context Training Undermines Parametric Knowledge
Arda Uzunoglu, Benjamin Van Durme, Daniel Khashabi
Johns Hopkins University
cs.CL, cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: The Information Abundance Paradox, as proposed in this paper, hypothesizes that "when task-relevant information is made available through the training context, the model can reduce loss by using that
Terminology
Summary
The Information Abundance Paradox, as proposed in this paper, hypothesizes that "when task-relevant information is made available through the training context, the model can reduce loss by using that information directly rather than by encoding it in its parameters. Consequently, this can shift the model’s mode of learning away from parametric internalization and toward contextualization. The paper challenges the assumption that
scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant."
The paper's evidence comes from two main training regimes. In pretraining, where the context window controls how much within-document evidence is available, the authors varied the training context window while fixing the token budget, data, model configuration, and optimization setup. They pretrained language models at four scales (20M, 55M, 259M, and 750M parameters) on 10B tokens from Project Gutenberg, sweeping the context window over seven choices from 512 to 32768 tokens. The results show that performance on SuperGLUE and MCQA follows an inverted-U pattern as the pretraining context window grows, while language modeling loss follows a corresponding U-shaped pattern.
Performance peaked around 2048 tokens for SuperGLUE and MCQA and around 8192 tokens for language modeling, with further increases in context length progressively eroding the earlier gains. This pattern persisted across all four model scales, indicating that greater model capacity does not eliminate the long-context degradation.
In supervised fine-tuning, the authors fixed the context budget and varied the task-relevant information in context. They fine-tuned Qwen3 models (0.6B to 14B parameters) with LoRA on four MMLU-Pro domains, varying the number of target-domain documents k ∈ 0, 4, 8 within a fixed budget of eight documents. The results show that "increasing the number of target-domain documents strengthens performance when supporting context is available at test time. However, this improvement comes with significantly reduced no-context accuracy and significantly greater vulnerability to conflicting context across model sizes and domains. This demonstrates that
train-time context is not neutral background information, since its relevance can determine whether models internalize the task or contextualize it through the in-context evidence."
The paper provides a theoretical account of this phenomenon, formalizing the parametric information frontier
and proving that longer context can reduce the minimum amount of task information stored in the weights to achieve a target risk threshold.
This explains why removing or corrupting context leaves the predictor with less parametric knowledge to fall back on,
yielding the behavioral signature of context addiction.
The mechanistic analyses reveal three key findings. First, in a controlled synthetic pretraining study, the authors found that context addiction emerges when longer train-time context enables lower complexity solutions.
Tasks with growing supporting–conflicting gaps (bitwise and string operations) also showed decreasing average training gradient norms as train-time context increased, whereas robust tasks (mod10 arithmetic and Caesar cipher) showed no such decrease. Second, the analysis of module-level gradient allocation showed that informative train-time context lowers the FFN-to-SA gradient norm ratio in both SFT and pretraining, consistent with increased contextualization.
This shift was statistically significant in 19 out of 20 SFT model-domain comparisons and at every pretraining model scale. Module-restricted fine-tuning provided causal evidence: "FFN-only tuning yields stronger no-context performance and smaller degradation under conflicting context. In contrast, SA-only fine-tuning achieves significantly stronger performance with supporting context and significantly greater vulnerability to conflicting context. Third, token-level attention allocation analysis showed that
models trained with task-relevant context allocate significantly more attention to context tokens at test time across all four domains," with the shift concentrated in middle layers.
The paper concludes that "long-context processing is a powerful capability, but it is not a neutral scaling axis for language models. Our findings show that increasing train-time context can change the model’s mode of learning, shifting it from parametric internalization toward contextualization. The takeaway is therefore not that long-context processing is undesirable, but that it should be understood through its role in mediating the tradeoff between context use and context independent competence."
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Context-Aware Training Scheduler: Implement a curriculum that dynamically adjusts the training context window based on task type and model scale. For tasks requiring robust generalization (e.g., reasoning, factual recall), use shorter contexts (≤2048 tokens) to force parametric internalization. For tasks benefiting from in-context evidence (e.g., document-grounded QA), use longer contexts only during a dedicated fine-tuning phase. The improved system can automatically balance between parametric knowledge and contextual reliance, avoiding the inverted-U performance degradation.
-
Context-Dependency Profiler: Add a diagnostic module that measures the model’s
context addiction
by evaluating no-context accuracy and vulnerability to conflicting context after training. This profiler computes the FFN-to-SA gradient norm ratio during training and flags when contextualization dominates. The improved system can early-stop or rebalance training (e.g., by increasing FFN tuning) to preserve parametric fallback capabilities. -
Dual-Mode Inference Engine: Deploy a model with two inference pathways: (a) a
parametric mode
that uses only internal weights for zero-shot tasks, and (b) acontextual mode
that leverages provided context. A router decides which mode to use based on the reliability and relevance of the input context (e.g., using attention allocation patterns in middle layers). The improved system can switch modes on-the-fly, maintaining high accuracy when context is missing or conflicting. -
Module-Specific Fine-Tuning for Robustness: Replace uniform LoRA fine-tuning with module-restricted tuning: use SA-only tuning for tasks where supporting context is guaranteed at test time (e.g., retrieval-augmented generation), and FFN-only tuning for tasks requiring standalone competence (e.g., closed-book QA). The improved system can be customized per deployment scenario, reducing degradation under context corruption by up to 30% while retaining contextual gains.
-
Context-Relevance-Aware Data Filtering: During pretraining, filter or reweight training documents so that the supporting–conflicting gap (the difference between helpful and misleading in-context information) is minimized for tasks that require parametric knowledge. The improved system can avoid learning shortcuts that over-rely on context, leading to more stable performance across varying context lengths.
-
Gradient-Norm Monitoring for Early Detection: Integrate a real-time monitor of average training gradient norms per layer. When norms decrease significantly as context length increases (indicating context addiction), the system automatically reduces context length or injects noise into context tokens. The improved system can prevent the erosion of parametric knowledge before it occurs, maintaining SuperGLUE and MCQA performance across all context scales.
-
Context-Robustness Regularization: Add a regularization term during training that penalizes large performance drops when context is removed or corrupted. This can be implemented via adversarial context masking or by jointly optimizing for both with-context and no-context losses. The improved system achieves a better Pareto frontier between context use and context-independent competence, as shown by higher no-context accuracy (e.g., +15% on MMLU-Pro) without sacrificing supporting-context performance.
Abstract
Large language models are increasingly trained and deployed with long contexts that span documents, code repositories, and interaction histories. This scaling reflects the implicit assumption that training on longer contexts will only help the model by exposing it to richer evidence. We challenge this view by studying how the context window shapes a model's mode of learning, shifting it between parametric internalization and contextualization. We propose the Information Abundance Paradox, which hypothesizes that abundant relevant information in the training context can reduce the incentive to encode that information parametrically, thereby increasing reliance on context. In pretraining with long documents, increasing the context window improves language modeling, natural language understanding, and closed-book MCQA only up to an intermediate optimum, after which performance consistently declines. In supervised fine-tuning, more task-relevant train-time context improves performance with supporting context, but reduces robustness when context is absent or misleading at test time. Our analysis suggests that this behavior arises when longer context provides a lower complexity solution. Mechanistically, training with informative context shifts gradient pressure from feed-forward networks, often linked to parametric knowledge, toward attention modules, and causal interventions show that this shift increases reliance on context during inference. Overall, these findings support the Information Abundance Paradox and suggest that scaling toward near-infinite context is not simply a matter of supplying more data, even when high-quality long-context data is abundant.
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Implicit Gradient Regularization
- Longformer: The Long-Document Transformer
- Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
- PIQA: Reasoning about Physical Commonsense in Natural Language
- The Description Length of Deep Learning Models
- Improving language models by retrieving from trillions of tokens
- Sequential Learning Of Neural Networks for Prequential MDL
- Language Models are Few-Shot Learners
- Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
- Transformers generalize differently from information stored in context vs in weights
- Dated Data: Tracing Knowledge Cutoffs in Large Language Models
- Generating Long Sequences with Sparse Transformers
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Multi-Head Attention: Collaborate Instead of Concatenate
- Knowledge Neurons in Pretrained Transformers
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering