Physically Partitioned KVCache Format for CPU--GPU Load Balancing in MoE Inference
cs.DC, cs.LG
Submitted: 2026-09-13
Updated: 2026-09-13
Code: https://github.com/ggerganov/llama.cpp
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- FastDecode: High-Throughput GPU-Efficient LLM Serving using Heterogeneous Pipelines
- Mixtral of Experts
- HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
- HeadInfer: Memory-Efficient LLM Inference by Head-wise Offloading
- Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving
- Qwen3 Technical Report
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- eLLM: Elastic Memory Management Framework for Efficient LLM Serving
- PowerInfer-2: Fast Large Language Model Inference on a Smartphone
- An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference
- ScoutAttention: Efficient KV Cache Offloading via Layer-Ahead CPU Pre-computation for LLM Inference
- Janus: Disaggregating Attention and Experts for Scalable MoE Inference
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing