vToken: Token-Level Virtualization for Reclaimable KV Caches
Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li
National University of Defense Technology · Peking University
cs.AI, cs.DC, cs.OS
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/anthropics/claudecode
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement in KV cache management for LLM serving systems.
Terminology
Summary
vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement in KV cache management for LLM serving systems. It addresses the granularity mismatch between token-level KV eviction algorithms (like H2O, StreamingLLM, Scissorhands, Random) and block-level memory management (PagedAttention in vLLM). This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable because blocks containing both evicted and retained tokens cannot be released.
vToken maintains a stable logical token view through a per-sequence token-table indirection and realizes physical reclamation by asynchronously repacking live tokens. It preserves PagedAttention kernels and CUDA Graph compatibility. The design includes a token table for logical-to-physical mappings, a reclamation backend with lazy compaction, and scheduler hooks for slot-mapping refresh and CUDA-event dependency.
Implemented in vLLM and evaluated with H2O, Random, and Scissorhands across models, vToken reduces retained KV blocks per request by 27.2%–72.3% and improves SLA-constrained throughput by up to 1.37× compared with a paired Naive-Evict baseline. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2×, while reducing the per-policy integration footprint from 500+ lines to under 50.
Key results include: memory utilization improvements of 21.88% on Llama-3.1-8B and 21.67% on Mistral-7B; SLA-constrained throughput improvements of 9.9%–37.3% on Mistral-7B and 18.9% average on Llama-3.1-8B; capacity frontier extension from C=5 to C=8 at gpu mem util=0.35 and from C=11 to C=22 at gpu mem util=0.50; and a 2× frontier extension on Qwen2.5-14B. The indirection-only ablation changes throughput and p95 by less than 1.0% relative to Native vLLM. Generation stability checks show a mean paired ROUGE-L F1 difference of only −0.0016, with 93.1% of pairs differing by at most 0.01.
Improvements for AI systems
Improvements to AI systems:
-
Unified token-eviction and memory-management layer: Integrate vToken’s token-table indirection and asynchronous repacking directly into LLM serving frameworks (e.g., vLLM, TensorRT-LLM) so that any token-level eviction policy (H2O, StreamingLLM, Scissorhands, Random) can run without modifying kernel code or CUDA Graphs. This reduces integration effort from 500+ lines to <50 lines per policy.
-
Fragmentation-free KV cache under aggressive eviction: Enable systems to reclaim memory from blocks that contain a mix of evicted and live tokens. The improved system can maintain a stable logical token view while physically compacting live tokens into fewer blocks, eliminating intra-block fragmentation. This directly increases effective KV memory utilization by 22% on Llama-3.1-8B and Mistral-7B.
-
Higher SLA-constrained throughput: By reducing retained KV blocks per request by 27.2%–72.3%, the improved system can serve more concurrent requests under the same latency SLO. It achieves up to 1.37× throughput improvement over naive eviction baselines, and 9.9%–37.3% improvement on Mistral-7B and 18.9% average on Llama-3.1-8B.
-
Extended concurrency and capacity frontier: With a constrained active-KV budget, the improved system can double the maximum feasible concurrency (e.g., from C=5 to C=8 at gpu mem util=0.35, and C=11 to C=22 at gpu mem util=0.50). This allows serving longer sequences or more users on the same GPU memory.
-
Policy-agnostic memory reclamation: The improved system can apply any eviction policy without sacrificing memory efficiency. It decouples logical liveness from physical placement, so even random eviction (which is highly fragmented) becomes memory-efficient, enabling simpler and faster eviction heuristics without performance loss.
-
Negligible overhead for non-eviction workloads: The indirection layer adds <1.0% throughput and p95 latency overhead compared to native vLLM when no eviction is used. Thus, the improved system can be deployed as a default memory manager, safe for all workloads.
-
Stable generation quality: The improved system preserves output quality, with mean paired ROUGE-L F1 difference of only −0.0016 and 93.1% of pairs differing by at most 0.01, ensuring that memory optimizations do not degrade text generation coherence.
What the improved AI system can do:
-
Serve LLMs with 2× higher concurrency on the same GPU memory, while maintaining SLOs.
-
Run any token-eviction algorithm (including random) with near-optimal memory utilization, without kernel changes.
-
Seamlessly switch between eviction and non-eviction modes with <1% overhead.
-
Support longer context windows or larger batch sizes under memory pressure, improving real-world serving economics and user experience.
Sources
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs
- DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
- GPT-4 Technical Report
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
- DeepSeek-V3 Technical Report
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
- Fast Transformer Decoding: One Write-Head is All You Need
- Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference
- Attention Is All You Need
- Efficient Streaming Language Models with Attention Sinks
- Strata: Hierarchical Context Caching for Long Context Language Model Serving
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection