Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster".
Dev: The electric power supply for AI datacenters has become a critical bottleneck in achieving Artificial General Intelligence, making end-to-end power management across planning, deployment validation, and runtime optimization essential.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So, building on that context, let’s talk about the actual summary of "Provisioning to Runtime Optimization of a one hundred MW-Scale AI Cluster <ref:2605.24461#pg0>." It outlines how they move from initial power planning through deployment validation and then into dynamic runtime tuning for this large cluster.
Dev: Essentially, the paper maps out three distinct phases: first, datacenter-planning where they figure out the best cluster configuration based on throughput versus power budget; second, deployment-validation where they analyze real data to adjust assumptions about GPU limits; and third, operational phase where they dynamically tune settings during live workloads.
Taro: The summary emphasizes that the initial planning objective is to maximize total throughput T(p) within a fixed power budget Ptotal by optimizing the performance-per-watt ratio η(p).
Rosa: That optimization leads them to a specific finding, which is that they concluded the optimal Perf/Watt operating point for their cluster occurs when setting the GPU power limit to approximately one thousand W <ref:2605.24461#pg1>.
Dev: That one thousand W setting seems central, and it’s derived from balancing performance reduction against increased GPU density within the constraints of Equation two which defines total datacenter power consumed per GPU.
Taro: When you look at the deployment validation phase, they refine these assumptions by analyzing actual power consumption data and workload performance to make adjustments to those power limits across the entire delivery hierarchy.
Rosa: A key result from that validation was determining that for their GB200 accelerators, the optimal setting for maximizing cluster performance was actually nine hundred sixty W, which is about eighty percent of the TDP <ref:2605.24461#pg1>.
Dev: That nine hundred sixty W figure is quite specific; it shows that empirical data significantly refined the initial theoretical models before they moved into live operation <ref:2605.24461#pg1>.
Taro: If we consider the implications of this validation step, it suggests that relying solely on vendor specifications for power limits is insufficient when dealing with real-world hardware mixes in a datacenter.
Rosa: It really underscores how important it is to use empirical data from actual power delivery hierarchy monitoring to set those operational margins accurately.
Dev: And that leads directly into the operational phase, which focuses on dynamically tuning settings during uncontrolled, live workloads to maximize performance within the power budget.
Taro: During runtime tuning, they employ several advanced techniques like temporal averaging based on moving-average power instead of instantaneous spikes to respect different time scales.
Rosa: That temporal averaging is clever because it acknowledges that an RPP might tolerate a ten percent overdraw for seventeen minutes but trip in sixty seconds at forty percent, which is very relevant for system safety.
Dev: And they also use spatial averaging to account for millisecond divergence across GPUs and racks, allowing them to provision per-rack power based on a lower percentile of the aggregate distribution.
Taro: The introduction of the Dimmer for training clusters is another advanced technique that smoothly reduces GPU power across all affected racks when the total draw approaches ninety-seven percent of its limit.
Rosa: So, this whole process shows a progression from static planning to real-time, adaptive control to handle the inherent variability in AI workloads and physical infrastructure.
The paper's summary: Dev: Now we’re looking at the suggested improvements for "Provisioning to Runtime Optimization of a one hundred MW-Scale AI Cluster," which are basically actionable steps to take based on their findings <ref:2605.24461#pg0>.
Rosa: One major suggestion is moving away from using the static Thermal Design Power directly; instead, they recommend implementing a predictive model that uses the accelerator’s projected power-performance curve to determine dynamic, lower power limits based on expected workload intensity.
Taro: That makes sense because it addresses the issue of heterogeneity across different workloads; if we know what the AI system is expected to do, we can set a more realistic constraint for its hardware.
Dev: Furthermore, in the planning phase they suggest adopting a configuration strategy that favors newer accelerators, like the GB200 over older ones like H100 because they offer better performance-per-watt.
Rosa: That aligns with their earlier findings that newer hardware often offers superior efficiency, even if it means fitting fewer physical GPUs into the same fixed power budget initially.
Taro: They also suggest modeling the power delivery hierarchy explicitly during planning to spot and mitigate bottlenecks caused by mixes of GPUs, networking gear, and CPUs across different racks before they become actual problems.
Dev: That proactive modeling addresses the problem mentioned earlier about hardware heterogeneity causing power imbalances across different delivery paths.
Rosa: In the deployment validation phase, they propose a validation loop that uses real-time data from Reactor Power Panel sensors to calibrate and adjust readings from the Power Supply Units more accurately.
Taro: Using the seventh percentile aggregation method for estimating rack power instead of just maximum PSU readings seems like a smarter way to account for variability in real-world conditions.
Dev: And they suggest using performance monitoring during validation to empirically refine those operational power-limit ranges based on what is actually measured under representative workloads.
Rosa: In the operational phase, they propose an "always-on" software power smoothing kernel that runs exclusively on register data to generate synthetic load, which helps maintain a stable power draw profile without heavy application overhead.
Taro: That synthetic load generation sounds like a very clean way to handle synchronized power oscillations during training, keeping the system stable while the main job executes.
Dev: And for runtime capping, they suggest a job scheduler-aware dynamic power capping mechanism, like a Dimmer that prioritizes critical pre-training jobs when power limits are approached.
Rosa: That scheduling awareness is important because it means you don't just blindly throttle everything; you can make intelligent decisions about what gets slowed down during peak stress.
Taro: I think the combination of these suggestions shows a comprehensive approach to managing the system from design through operation, addressing both static constraints and dynamic execution issues.
The paper's improvements: Dev: So, to wrap up this discussion on "Provisioning to Runtime Optimization of a one hundred MW-Scale AI Cluster," we’ve covered how they systematically move from high-level planning down to the specific runtime tuning techniques they employ <ref:2605.24461#pg0>.
Rosa: The overall implication is that for deploying large-scale AI infrastructure, we need an end-to-end power management experience that addresses planning, validation, and dynamic optimization simultaneously.
Taro: The work highlights that hardware heterogeneity across racks creates constraints in power delivery paths that require a systemic view rather than just local fixes to solve them.
Dev: They demonstrated how optimizing the Perf/Watt ratio to an empirical sweet spot of around one thousand W for GB200s provides better overall efficiency and throughput compared to simply maximizing per-GPU performance <ref:2605.24461#pg1>.
Rosa: It’s about finding that balance, realizing that a slightly reduced power setting often leads to a measurable improvement in total cluster throughput.
Taro: For autonomous systems, this suggests that robust autonomy needs to account for these deep physical constraints on power delivery infrastructure when designing agents for complex environments.
Dev: We're looking at how their methodology informs our understanding of loop rates and the latency implications when dealing with these dynamic power adjustments under load.
Rosa: Ultimately, "Provisioning to Runtime Optimization of a one hundred MW-Scale AI Cluster" provides a detailed roadmap for engineers trying to build systems that operate safely and efficiently within tight electrical envelopes <ref:2605.24461#pg0>.
Conclusion: Rosa: So we’ve seen how researchers tackle the massive power demands of modern AI data centers in this paper, "Provisioning to Runtime Optimization of a one hundred MW-Scale AI Cluster."
Dev: Exactly, and what really stands out is the meticulous three-phase approach they take: planning, validation, and then dynamic runtime tuning.
Taro: I think the most fascinating part for autonomy applications is how they handle that spatial averaging during runtime to manage millisecond divergence across racks.
Rosa: That makes total sense from a field perspective; if a system can’t account for those tiny, fast fluctuations in power draw across different components, it’s going to cause real instability when things are running uncontrolled.
Dev: I agree with Rosa on the instability; the paper shows that relying on instantaneous power readings is just not viable for maintaining safe loop rates and failure modes under live workloads.
Taro: And from an autonomy standpoint, if a robot needs to operate in an unpredictable environment, having a framework that smooths out those transient spikes during intensive processing is crucial for reliable decision-making.
Rosa: It does feel like this level of infrastructure control might be necessary before we can even talk about deploying complex AI systems outside of controlled labs for extended periods.
Dev: I’m curious though, Rosa, how long do you think these dynamic adjustments could hold up in a truly uncontrolled environment where the power delivery itself is less predictable?
Rosa: That’s a tough one, Dev; the paper focuses heavily on lab-scale validation and controlled environments for those temporal averaging techniques.
Taro: But if we apply that concept to field robotics, maybe we can build predictive models that anticipate power draw based on the robot's immediate task intensity?
Dev: That brings us right back to the planning phase, Taro; getting that predictive model right is key before you even worry about runtime adjustments.
Rosa: Well, for now, this paper gives us a solid blueprint for how to manage the power side of these massive AI deployments.
Dev: It really shows that optimizing performance-per-watt isn't just about picking the best chip; it’s about mastering the entire power delivery hierarchy from start to finish.
Taro: I think their work on balancing throughput against power budget sets a good precedent for how we should approach resource allocation in any large, complex system.
Rosa: Absolutely, and it makes me wonder what happens when we start scaling these concepts to even larger infrastructures or different types of processing units.
Meta Platforms
cs.AR, cs.DC, cs.SY, eess.SY
Submitted: 2026-05-23
Updated: 2026-10-02
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: The electric power supply for AI datacenters has become a critical bottleneck in achieving Artificial General Intelligence, making end-to-end power management across planning, deployment validation,
Key concepts
- Datacenter-Planning Phase
- This initial stage determines the best cluster configuration by solving for maximum throughput given a total power limit. The goal is to find the optimal number of accelerators and their power limits that maximize performance while staying within the fixed 150 MW budget.
- Performance-per-Watt Ratio (Perf/Watt)
- This metric measures how efficiently a GPU performs relative to the power it consumes. The research found that operating at a specific Perf/Watt point, around 1000W per GPU, balances performance gains against density limitations for the cluster.
- Temporal Averaging
- This technique manages power limits by using moving averages instead of instantaneous power spikes. This prevents sudden trips while respecting time scales where a system might tolerate temporary overdraws, ensuring smoother operation during live workloads.
- Dimmer for Training Clusters
- This is an operational strategy that gradually reduces GPU power across all racks simultaneously when the total cluster draw nears its limit. It smooths out transient spikes and ensures continuous job execution even under high load.
Terminology
Summary
The electric power supply for AI datacenters has become a critical bottleneck in achieving Artificial General Intelligence, making end-to-end power management across planning, deployment validation, and runtime optimization essential. This work presents an end-to-end production experience for a 150 MW cluster of 83K GB200 GPUs, detailing strategies to maximize compute capacity within fixed power budgets by optimizing the performance-per-watt ratio across the entire power delivery hierarchy.
The gist
This paper describes the end-to-end power management process for a hyperscale AI datacenter, covering early planning, deployment validation, and dynamic runtime tuning for a 150 MW cluster of 83K GB200 GPUs.
How it works: Three Phases of Power Management
The paper outlines three distinct phases of power management that are addressed in the research: (1) datacenter-planning phase, (2) deployment-validation phase, and (3) operational phase. The goal is to manage power from initial planning through dynamic tuning during live workloads.
(1) Datacenter-Planning Phase:
The primary objective here is to determine the optimal cluster configuration that maximizes overall throughput (T(p)) at the datacenter level within the fixed power budget (Ptotal).
This involves selecting the total number of accelerators, which dictates the permissible power limit per unit, and then solving for:
max p T(p) = N(p)· f(p) (1)
The methodology focuses on maximizing performance-per-watt ratio η(p) = f(p)/g(p),
where g(p), the total datacenter power consumed per GPU, is defined by Equation 2. The paper notes that this optimization leads to the conclusion that the optimal Perf/Watt operating point for the cluster is achieved when GPUs power limit is set to approximately 1000 W,
as this balances performance reduction with increased GPU density.
(2) Deployment-Validation Phase:
During this phase, operators refine assumptions by analyzing real-world data. This involves analyzing power consumption data and workload performance to make informed adjustments to GPU power limits and margins across power delivery hierarchy, ensuring they align with real-world conditions.
A key finding is that the optimal setting for the GB200 was determined to be a 960 W (80% of the TDP) as the Perf/Watt optimized setting to maximize the cluster performance.
(3) Operational Phase:
The operational phase focuses on dynamically tuning power settings during the execution of uncontrolled, live workloads to maximize performance within the power budget.
This involves several advanced techniques:
-
Temporal averaging is used to enforce limits based on
moving-average power rather than instantaneous spikes,
respecting time scales where an RPP may tolerate a 10% overdraw for 17 minutes but trip in 60 seconds at 40%. -
Spatial averaging accounts for
millisecond divergence
across GPUs and racks, allowing per-rack power to be provisioned closer to a lower percentile of the aggregate distribution. -
The paper introduces
Dimmer for training clusters,
an approach that adopts agradually reduces GPU power across all racks within the affected power device
when total draw approaches 97% of its limit, ensuring jobs continue executing while smoothing transient spikes.
Key Insights and Findings
The research yielded several critical insights regarding system constraints and optimization:
-
In the datacenter-planning phase, the paper does not use the accelerator’s Thermal Design Power (TDP) directly to determine capacity; instead, it assumes a
lower power limit based on the accelerator’s projected power-performance curve.
-
Hardware heterogeneity causes
considerable power imbalances across different power delivery paths,
meaning increasing limits for all accelerators is constrained by the path with the least headroom. -
Accelerator performance does not scale uniformly; for instance, GB200’s HBM bandwidth
stays constant as the power limit decreases from 1200 Watt (W) to 1000 W, but drops sharply by 15% when further reduced to 800 W.
-
In large-scale synchronous AI training, local management is ineffective; instead,
when a power-delivery device nears its limit, we lower the power limits for all accelerators it governs.
-
The paper found that operating at a slightly reduced power (1000W) yields better overall efficiency and throughput than maximizing per-GPU performance (1200W), resulting in a
9% improvement
in total cluster throughput compared to the 1200W baseline.
Improvements for AI systems
Here are specific improvements for AI systems based on this paper, categorized by the phase of power management they address:
)Phase 1: Planning and Provisioning Improvements (Pre-Deployment)
-
Instead of using the static Thermal Design Power (TDP), implement a predictive model that uses the accelerator’s projected power-performance curve to determine a dynamic, lower power limit for each accelerator based on its expected workload intensity.
-
In datacenter planning, adopt a configuration strategy that selects the latest accelerators (like GB200) over older ones (like H100) because they offer superior performance-per-watt, even if it means slightly fewer physical GPUs fit within the fixed power budget.
-
Model the power delivery hierarchy explicitly during planning to identify and mitigate bottlenecks caused by heterogeneous hardware mixes across different racks (GPU, Network, CPU). This allows for pre-emptive allocation of resources to avoid future communication or power delivery constraints.
)Phase 2: Deployment Validation Improvements (Refinement)
-
Implement a validation loop that uses real-time data from Reactor Power Panel (RPP) sensors to calibrate and adjust the readings from the Power Supply Units (PSU), specifically by adopting the 70th percentile (P70) aggregation method for rack power estimation, rather than relying on maximum PSU readings.
-
Use performance monitoring during validation to refine the operational power-limit ranges for accelerators based on measured consumption under representative workloads, ensuring these limits are empirically sound rather than just vendor specifications.
)Phase 3: Active Operation Phase Improvements (Runtime Optimization)
-
Implement an
always-on
software power smoothing kernel that runs exclusively on register data to generate synthetic load, effectively mitigating synchronized power oscillations during training. This maintains a stable, high-power draw profile without significant application overhead. -
Develop a job scheduler-aware dynamic power capping mechanism (Dimmer) that prioritizes mission-critical pre-training jobs over lower-priority ones when power limits are approached. This ensures that transient spikes do not disproportionately throttle large, synchronous training jobs, maintaining overall cluster throughput.
-
Dynamically adjust the per-GPU power limit during live workloads using a seven-second averaging window to detect sustained overages versus brief surges, ensuring safety while maximizing compute utilization (e.g., moving from a conservative 960W baseline to an optimized 1020W limit when necessary).
)Overall System Capability of the Improved AI System
The improved system will be a hyperscale AI training cluster capable of:
-
Maximizing aggregate throughput by dynamically balancing the trade-off between per-GPU performance and overall GPU density, achieving up to a 10% increase in throughput compared to static provisioning.
-
Operating safely within strict power budgets by accurately modeling and respecting the complex, hierarchical constraints of the electrical infrastructure (MSB/SB/RPP).
-
Maintaining high stability during large-scale synchronous training jobs by proactively smoothing power oscillations, preventing straggler effects caused by transient power spikes, and intelligently managing resource contention across different job priorities.
-
Achieving higher energy efficiency (Perf/Watt) by optimizing the per-GPU power limit to an empirically determined sweet spot (e.g., 1000W for GB200), leading to better utilization of the total power envelope before physical scaling is limited by infrastructure constraints.
Abstract
The electric power supply for AI data centers is now the most significant bottleneck in the race toward Artificial General Intelligence, surpassing even the constraint of AI accelerator availability. To our knowledge, this paper is the first to describe the end-to-end power management process for a hyper-scale AI datacenter; from early power planning to accommodate next-generation accelerators 6--12 months before their general availability, to tuning power settings after large scale deployment, and finally to dynamic, runtime power management for evolving workloads. We present detailed power measurements for a 150 MW datacenter hosting a cluster of 83K GB200 GPUs. We share insights from building this state-of-the-art AI cluster. We hope this work encourages practitioners across the industry to share their own experiences as well.
Sources
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4
- SISA: A Scale-In Systolic Array for GEMM Acceleration