PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response
cs.DC, cs.AI
Submitted: 2026-08-22
Updated: 2026-08-22
Code: https://github.com/Azure/AzurePublicDataset
Terminology
Sources
- SLOs-Serve: Optimized Serving of Multi-SLO LLMs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- EcoServe: Designing Carbon-Aware AI Inference Systems
- Splitwise: Efficient generative LLM inference using phase splitting
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
- Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
- VoltanaLLM: Energy-Efficient and SLO-Aware Disaggregated LLM Serving via Adaptive Frequency Control and State-Space Routing
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
- PolyServe: Efficient Multi-SLO Serving at Scale
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing