Where Should a Document Live: Context, Representations, or Parameters?
cs.CL, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
Code: https://github.com/uhh-hcds/g4kmu-paper
Project page: https://nyu-mll.github.io/quality
License: http://creativecommons.org/licenses/by/4.0/
The gist: To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the
Terminology
Abstract
To answer questions outside of their pre-training data, large language models (LLMs) need access to new information, which can be presented in the context window as documents, encoded into the model's parameters, or injected as latent representations. However, each of these methods comes with different efficiency, cost, and performance trade-offs, with no single winner. We present a controlled comparison of representation-based (KV-cache based) and parametric (fine-tuning-based) adaptation methods on five knowledge-intensive benchmarks. We show that in the oracle setting, Cartridges (KV) are the most accurate injection method at nearly every storage budget, outperforming parametric methods by 10 points. Compaction (KV) matches Cartridges only at low compression rates, lagging behind the parametric methods by 10 points at rates higher than 50 times. In the more realistic multi-document retrieval scenario, Cartridges are the only method that matches in-context learning (ICL), leading the parametric methods by 29 points and Compaction by 15 points. Nonetheless, Cartridges are also the only method, besides full fine-tuning and large MLP adapters, that suffers from catastrophic forgetting, i.e., a 6% performance degradation on control benchmarks, with 13% in coding.
Sources
- Training Plug-n-Play Knowledge Modules with Deep Context Distillation
- KV Packet: Recomputation-Free Context-Independent KV Caching for LLMs
- Punica: Multi-Tenant LoRA Serving
- Evaluating Large Language Models Trained on Code
- LongHealth: A Question Answering Benchmark with Long Clinical Documents
- gpt-oss-120b & gpt-oss-20b Model Card
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- C$^2$KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Cartridges at Scale: Training Modular KV Caches over Large Document Collections
- Multi-LoRA Composition for Image Generation
- Instruction-Following Evaluation for Large Language Models
- Fast KV Compaction via Attention Matching
- Understanding Parametric Knowledge Injection in Retrieval-Augmented Generation
- Gemma 3 Technical Report
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering