Calibrate, Then Route: A Measured Study of Learned Request Routing for Disaggregated LLM Serving
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools.
Terminology
Abstract
Disaggregated LLM serving places compute heavy prefill and memory heavy decode on separate GPU pools. Systems such as DistServe, Splitwise, and Mooncake make this separation fast, but routing still determines which instances handle each request. We study a router that estimates the additional completion time on each instance using exact prompt length, predicted output length, post admission KV cache pressure, and SLO class. We develop the policy in a discrete event simulator and validate it on eight NVIDIA A40 GPUs, each running a vLLM engine, with NIXL transferring KV caches between pools. All workloads run at measured saturation. Across three mixed, bursty arrival traces, the calibrated router achieves the highest mean goodput at 0.864, compared with 0.835 to 0.847 for round robin, least loaded, and a length heuristic. It also shows the lowest variance across traces. It beats round robin and the length heuristic on all three traces and least loaded on two. On the third, it trails by 0.003, within run to run noise. Hardware calibration matters: simulator derived constants cost 4.5 goodput points and roughly 40 percent of the tail latency advantage, reducing the scorer to little more than queue counting. Benefits grow with decode pool size and traffic heterogeneity but disappear in pools with three instances, where queue counts are often enough. Under extreme scarcity, greedy cost minimization concentrates requests on the cheapest scored instance, and blind spreading performs better. With calibrated costs, the learned router matches the goodput of round robin using six GPUs instead of seven.
Sources
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- Splitwise: Efficient generative LLM inference using phase splitting
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- Llumnix: Dynamic Scheduling for Large Language Model Serving
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference Pipeline
- Efficient LLM Scheduling by Learning to Rank
- Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
- LAPS: A Length-Aware-Prefill LLM Serving System
- STAR: Decode-Phase Rescheduling for LLM Inference
- LLM Inference Serving: Survey of Recent Advances and Opportunities
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- H$_2$O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
- Vidur: A Large-Scale Simulation Framework For LLM Inference
- The Llama 3 Herd of Models
- Qwen2.5 Technical Report
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
- Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection