Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
cs.CL
Submitted: 2026-09-03
Updated: 2026-09-28
Code: https://github.com/SalesforceAIResearch/Random-Attention
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck.
Terminology
Abstract
Large language models achieve superior performance on tasks that require extended reasoning, but long chains of thought make the KV cache a severe memory bottleneck. Existing KV cache compression methods share one paradigm: score each cached token by some estimate of how much it will matter later, and keep the top-scoring ones. We show that the selection signal contributes almost nothing. Random Attention keeps the prompt and evicts uniformly at random within each attention head, computing no score at all; across four models and six reasoning tasks it matches the strongest prior evictor while serving 32-43% higher throughput than it in vLLM deployment. Controlled experiments explain this by showing that 1) the prompt is the fragile part of the cache, and most of the gap between selectors is just whether their selection signal happened to keep it; 2) the reasoning trace protects itself against eviction with redundancy at two levels, in the text (the model restates what it still needs as it works) and across attention heads (each keeps its own copy of the trace), so once the prompt is safe, a random draw retains enough copies of what the model still needs, and no score is required to pick them. Our code is publicly available at https://github.com/SalesforceAIResearch/Random-Attention.
Sources
- Phi-4-reasoning Technical Report
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Value-Aware Stochastic KV Cache Eviction for Reasoning Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Protection Is (Nearly) All You Need: Structural Protection Dominates Scoring in Globally Capped KV Eviction
- Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
- Prefix Sliding for efficient test-time scaling
- Coverage-Driven KV Cache Eviction for Efficient and Improved Inference of LLM
- How Much Cache Does Reasoning Need? Depth-Cache Tradeoffs in KV-Compressed Transformers
- Qwen3 Technical Report
- Not All Thoughts Need HBM: Semantics-Aware Memory Hierarchy for LLM Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering