DLB: Distributed Load Balancing at Scale for Generative AI Inference
cs.DC, cs.SY, eess.SY
Submitted: 2026-09-17
Updated: 2026-09-17
Code: https://github.com/vllm-project/aibrix
Terminology
Sources
- Load Balancing with Network Latencies via Distributed Gradient Descent
- GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving
- SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM Inference
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM Serving
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing