KV Cache Offloading for Context-Intensive Tasks
cs.LG, cs.AI, cs.CL
Submitted: 2026-04-09
Updated: 2026-09-01
Comments: Preprint
Code: https://github.com/ByteDance-Seed/ShadowKV
Project page: https://nvidia.github.io/TensorRT-LLM/features/quantization.html
License: http://creativecommons.org/licenses/by/4.0/
The gist: With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage.
Terminology
Abstract
With the growing demand for long-context LLMs across a wide range of applications, the key-value (KV) cache has become a critical bottleneck for both latency and memory usage. Recently, KV-cache offloading has emerged as a promising approach to reduce memory footprint and inference latency while preserving accuracy. Prior evaluations have largely focused on tasks that do not require extracting large amounts of information from the context. In this work, we study KV-cache offloading on context-intensive tasks: problems where the solution requires looking up a lot of information from the input prompt. We create and release the Text2JSON benchmark, a highly context-intensive task that requires extracting structured knowledge from raw text. We evaluate modern KV offloading on Text2JSON and other context-intensive tasks and find significant performance degradation on both Llama 3 and Qwen 3 models. Our analysis identifies two key reasons for poor accuracy: low-rank projection of keys and unreliable landmarks, and proposes a simpler alternative strategy that significantly improves accuracy across multiple LLM families and benchmarks. These findings highlight the need for a comprehensive and rigorous evaluation of long-context compression techniques.
Sources
- RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems
- Measuring AI Ability to Complete Long Software Tasks
- Cache Me If You Must: Adaptive Key-Value Quantization for Large Language Models
- SnapStream: Efficient Long Sequence Decoding on Dataflow Accelerators
- The Pitfalls of KV Cache Compression
- Understanding the Physics of Key-Value Cache Compression for LLMs through Attention Dynamics
- KVSharer: Efficient Inference via Layer-Wise Dissimilar KV Cache Sharing
- RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval
- Transformer Acceleration with Dynamic Sparse Attention
- LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
- SCBench: A KV Cache-Centric Analysis of Long-Context Methods
- Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- KVzap: Fast, Adaptive, and Faithful KV Cache Pruning
- SpargeAttention: Accurate and Training-free Sparse Attention Accelerating Any Model Inference
- Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
- The Llama 3 Herd of Models
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
- LLaMA: Open and Efficient Foundation Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks