Where Should the KV Cache Live? Placement Policies Across GPU, CPU, and SSD for Long-Lived Sessions
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state.
Terminology
Abstract
GPU high bandwidth memory is scarce and expensive, and KV caches consume much of it as chats, agent loops, and document question answering accumulate state. Systems such as Mooncake, LMCache, FlexGen, InfiniGen, and AttentionStore extend GPU memory with CPU DRAM and SSD. The harder question is which blocks belong in each tier, when to move or evict them, and whether prefetching helps. We study these choices in a discrete event simulator spanning GPU HBM, CPU DRAM, and SSD, calibrated against a random forest execution time predictor. We compare recency, reuse frequency, predicted reuse, and an EWMA predictor with prefetch lookahead across chat, agent, and document question answering workloads. Tiering supports 73.02 times more concurrent sessions per GPU and lowers cost per session by 62.04 times. These gains come from tier capacities of 1 plus 8 plus 64, not placement policy. Decode is compute bound at batch size one in our setup, so placement barely affects throughput. It mainly changes PCIe migration traffic and time to first token. Recency produces 2.30 times less migration traffic than reuse frequency for chat. Reuse frequency performs best for agents and document question answering. The existing predicted reuse policy is byte identical to recency, making its agent recommendation effectively recency. A genuine EWMA predictor changes behavior but still ranks behind reuse frequency on the workloads prediction was expected to help. Prefetching does not justify its bandwidth cost. Across the policy and cache size grid, even an oracle with knowledge of future requests never beats no prefetch on migration traffic. Workload specific placement can reduce data movement, but the predicted reuse and prefetch recommendations are not supported as implemented.
Sources
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
- InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- SGLang: Efficient Execution of Structured Language Model Programs
- ServerlessLLM: Low-Latency Serverless Inference for Large Language Models
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization
- H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- Splitwise: Efficient generative LLM inference using phase splitting
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- Keep the Cost Down: A Review on Methods to Optimize LLM' s KV-Cache Consumption
- Efficient Streaming Language Models with Attention Sinks
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection