CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference.
Terminology
Abstract
Despite their strong performance, large language models (LLMs) are bottlenecked by KV cache memory traffic during long-context inference. Sparse attention is widely used to accelerate LLM inference by computing exact attention over a selected subset of tokens. To recover the contribution of tokens excluded from exact attention, recent methods apply coarse-grained compensation to the omitted attention tail. However, existing methods typically select tokens based on attention mass and only then compensate for the unselected tokens. This decoupled design overlooks their interaction: selection should prioritize tokens that would leave the largest compensation error if omitted. To address this limitation, we introduce CompKV, the first compensation-aware sparse attention framework that divides tokens into blocks and explicitly optimizes selection for the downstream compensation mechanism. Our theoretical analysis shows that the residual left by block-level mean compensation is governed by both block attention mass and within-block logit variation. We approximate this residual using compact block-level statistics, yielding a deployable selection criterion. We further develop an efficient asynchronous implementation. Experiments on RULER and LongBench-Pro show that CompKV performs best among the evaluated sparse baselines while delivering up to a 6.85 times self-attention speedup over full attention.
Sources
- LongBench Pro: A More Realistic and Comprehensive Bilingual Long-Context Evaluation Benchmark
- UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
- The Llama 3 Herd of Models
- GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks