The KV Cache Is the New Memory Wall
cs.DC, cs.LG, cs.PF
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/LMCache/LMCache
Terminology
Sources
- The Llama 3 Herd of Models
- Optimizing AI Inference Across the Deployment Stack
- Fast Transformer Decoding: One Write-Head is All You Need
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- KV Cache Optimization Strategies for Scalable and Efficient LLM Inference
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Mixtral of Experts
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- AI Benchmark Democratization and Carpentry
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing