PipeLive: Efficient Live In-place Pipeline Parallelism Reconfiguration for Dynamic LLM Serving
cs.DC, cs.LG
Submitted: 2026-04-14
Updated: 2026-09-23
Terminology
Sources
- The Llama 3 Herd of Models
- M'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless Clusters
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public Clouds
- GPT-4 Technical Report
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- LLaMA: Open and Efficient Foundation Language Models
- Towards Efficient and Practical GPU Multitasking in the Era of LLM
- Qwen3 Technical Report
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing