Analytical Power-Aware Provisioning for Prefill-Decode Disaggregated AI Inference
cs.DC, cs.PF, cs.SY, eess.SY
Submitted: 2026-09-21
Updated: 2026-09-21
Code: https://github.com/deepseekai/open-infra-index
Terminology
Sources
- Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
- DualScale: Energy-Efficient Disaggregated LLM Serving via Phase-Aware Placement and DVFS
- FleetOpt: Analytical Fleet Provisioning for LLM Inference with Compress-and-Route as Implementation Mechanism
- inference-fleet-sim: A Queueing-Theory-Grounded Fleet Capacity Planner for LLM Inference
- Electricity Demand and Grid Impacts of AI Data Centers: Challenges and Prospects
- From Tokens to Energy Flexibility: Quantization-Enabled Demand Response for Data Centers with LLM Inference Workloads
- Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
- WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs
- Frontier: Towards Comprehensive and Accurate LLM Inference Simulation
- PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
- BestServe: Serving Strategies with Optimal Goodput in Collocation and Disaggregation Architectures
- Power Aware Dynamic Reallocation For Inference
- Mitigation of Datacenter Demand Ramping and Fluctuation using Hybrid ESS and Supercapacitor
- A Theory of Probabilistic Power Provisioning for Data Centers with Distributed Energy Storage
- TokenScale: Timely and Accurate Autoscaling for Disaggregated LLM Serving with Token Velocity
- SLO-Aware Compute Resource Allocation for Prefill-Decode Disaggregated LLM Inference
- Taming the Chaos: Coordinated Autoscaling for Heterogeneous and Disaggregated LLM Inference
- PowerSlider: Exploiting Phase Asymmetry for LLM Serving under Demand Response
- GPU-to-Grid: Voltage Regulation via GPU Utilization Control
- Agentic Coding in the Wild: Characterizing GitHub Copilot Traces at Production Scale
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing