Compression Beyond the Uncompressed: A Two-Stage Training Recipe for Soft Context Compression in RAG
cs.CL
Submitted: 2026-09-04
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
The gist: Retrieval-Augmented Generation (RAG) enhances language models with external knowledge, but the lengthy retrieved context inflates the input and degrades inference efficiency.
Terminology
Abstract
Retrieval-Augmented Generation (RAG) enhances language models with external knowledge, but the lengthy retrieved context inflates the input and degrades inference efficiency. Soft context compression encodes each document into a substantially shorter embedding sequence. However, most existing approaches are trained by distilling outputs from uncompressed RAG systems, inherently limiting their performance relative to the original model. To address this limitation, we propose DEX-Comp, a two-stage training recipe: Pure Distillation warm-starts the compression model on the uncompressed RAG's correct responses only, and Hard Exploration then runs reinforcement learning solely on queries the uncompressed RAG fails, forcing the model to explore computation patterns better suited to compressed representations. On five open-domain QA benchmarks at retrieval depths from top-5 to top-30, DEX-Comp compresses retrieved contexts by 16 times and accelerates inference by 4 times -- 24 times, while achieving performance comparable to or exceeding the uncompressed RAG baseline across retrieval depths. Ablations and evaluations across diverse datasets and backbones further confirm the contribution of each stage and the generalization of our approach.
Sources
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Context Embeddings for Efficient Answer Generation in RAG
- ResidualKV: Residual-Based KV Cache Compression for Efficient Long-Context Inference
- Prompt Compression for Large Language Models: A Survey
- AttentionRAG: Attention-Guided Context Pruning in Retrieval-Augmented Generation
- Unlocking Context Constraints of LLMs: Enhancing Context Efficiency of LLMs with Self-Information-Based Content Filtering
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
- OSCAR: Online Soft Compression And Reranking
- Less Is More: Elevating RAG via Performance-Driven Context Compression
- Learning to Filter Context for Retrieval-Augmented Generation
- Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
- Autoencoding-Free Context Compression for LLMs via Contextual Semantic Anchors
- No Mean Feat: Simple, Strong Baselines for Context Compression
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Mistral 7B
- SPLADE-v3: New baselines for SPLADE
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering