vToken: Token-Level Virtualization for Reclaimable KV Caches

arXiv:2608.13263 · cs.AI, cs.DC, cs.OS · Submitted 2026-08-13 · Read on arXiv

Yuanhang Gao, Xiangrui Yang, Yuanfeng Chen, Hongjia Chen, Qianru Lv, Wenfei Wu, Dongsheng Li

National University of Defense Technology · Peking University

cs.AI, cs.DC, cs.OS

Submitted: 2026-08-13

Updated: 2026-08-14

Code: https://github.com/anthropics/claudecode

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement in KV cache management for LLM serving systems.

Terminology

Summary

vToken is a lightweight token-level virtualization layer that decouples logical token liveness from physical block placement in KV cache management for LLM serving systems. It addresses the granularity mismatch between token-level KV eviction algorithms (like H2O, StreamingLLM, Scissorhands, Random) and block-level memory management (PagedAttention in vLLM). This mismatch causes intra-block fragmentation, leaving a large fraction of allocated KV memory unreclaimable because blocks containing both evicted and retained tokens cannot be released.

vToken maintains a stable logical token view through a per-sequence token-table indirection and realizes physical reclamation by asynchronously repacking live tokens. It preserves PagedAttention kernels and CUDA Graph compatibility. The design includes a token table for logical-to-physical mappings, a reclamation backend with lazy compaction, and scheduler hooks for slot-mapping refresh and CUDA-event dependency.

Implemented in vLLM and evaluated with H2O, Random, and Scissorhands across models, vToken reduces retained KV blocks per request by 27.2%–72.3% and improves SLA-constrained throughput by up to 1.37× compared with a paired Naive-Evict baseline. Under a constrained active-KV budget, it extends the maximum feasible concurrency by up to 2×, while reducing the per-policy integration footprint from 500+ lines to under 50.

Key results include: memory utilization improvements of 21.88% on Llama-3.1-8B and 21.67% on Mistral-7B; SLA-constrained throughput improvements of 9.9%–37.3% on Mistral-7B and 18.9% average on Llama-3.1-8B; capacity frontier extension from C=5 to C=8 at gpu mem util=0.35 and from C=11 to C=22 at gpu mem util=0.50; and a 2× frontier extension on Qwen2.5-14B. The indirection-only ablation changes throughput and p95 by less than 1.0% relative to Native vLLM. Generation stability checks show a mean paired ROUGE-L F1 difference of only −0.0016, with 93.1% of pairs differing by at most 0.01.

Improvements for AI systems

Improvements to AI systems:

  1. Unified token-eviction and memory-management layer: Integrate vToken’s token-table indirection and asynchronous repacking directly into LLM serving frameworks (e.g., vLLM, TensorRT-LLM) so that any token-level eviction policy (H2O, StreamingLLM, Scissorhands, Random) can run without modifying kernel code or CUDA Graphs. This reduces integration effort from 500+ lines to <50 lines per policy.

  2. Fragmentation-free KV cache under aggressive eviction: Enable systems to reclaim memory from blocks that contain a mix of evicted and live tokens. The improved system can maintain a stable logical token view while physically compacting live tokens into fewer blocks, eliminating intra-block fragmentation. This directly increases effective KV memory utilization by 22% on Llama-3.1-8B and Mistral-7B.

  3. Higher SLA-constrained throughput: By reducing retained KV blocks per request by 27.2%–72.3%, the improved system can serve more concurrent requests under the same latency SLO. It achieves up to 1.37× throughput improvement over naive eviction baselines, and 9.9%–37.3% improvement on Mistral-7B and 18.9% average on Llama-3.1-8B.

  4. Extended concurrency and capacity frontier: With a constrained active-KV budget, the improved system can double the maximum feasible concurrency (e.g., from C=5 to C=8 at gpu mem util=0.35, and C=11 to C=22 at gpu mem util=0.50). This allows serving longer sequences or more users on the same GPU memory.

  5. Policy-agnostic memory reclamation: The improved system can apply any eviction policy without sacrificing memory efficiency. It decouples logical liveness from physical placement, so even random eviction (which is highly fragmented) becomes memory-efficient, enabling simpler and faster eviction heuristics without performance loss.

  6. Negligible overhead for non-eviction workloads: The indirection layer adds <1.0% throughput and p95 latency overhead compared to native vLLM when no eviction is used. Thus, the improved system can be deployed as a default memory manager, safe for all workloads.

  7. Stable generation quality: The improved system preserves output quality, with mean paired ROUGE-L F1 difference of only −0.0016 and 93.1% of pairs differing by at most 0.01, ensuring that memory optimizations do not degrade text generation coherence.

What the improved AI system can do:

  • Serve LLMs with 2× higher concurrency on the same GPU memory, while maintaining SLOs.

  • Run any token-eviction algorithm (including random) with near-optimal memory utilization, without kernel changes.

  • Seamlessly switch between eviction and non-eviction modes with <1% overhead.

  • Support longer context windows or larger batch sizes under memory pressure, improving real-world serving economics and user experience.

Sources

Related papers