PEEK: Predictive Queue-Informed KV Cache Management for LLM Serving
cs.DC, cs.AI
Submitted: 2026-05-10
Updated: 2026-09-12
Comments: 26 pages, 21 figures, 21 tables. Preprint, under review
Code: https://github.com/xiexbing/peek
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Locality-aware Fair Scheduling in LLM Serving
- RAGCache: Efficient Knowledge Caching for Retrieval-Augmented Generation
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge Fusion
- ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
- BatchLLM: Optimizing Large Batched LLM Inference with Global Prefix Sharing and Throughput-oriented Token Batching
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing