Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet".
Jane: The paper was written by Eliseo Curcio from Advanced Department of Artificial Intelligence and Energy – New York.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making the rounds, and the title alone tells you it's ambitious: "Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet." Jane, when you first saw that title, what went through your head?
Jane: Honestly, Tom, I thought it was two papers stapled together. One part is about controlling power on a single GPU, and the other is about an entire fleet of them. But the word "measured" is what caught my eye. They didn't just simulate this — they actually ran it on hardware.
Tom: Exactly. And that's the thing that separates this from a lot of the theoretical work we see. They took real A100 GPUs, ran actual GRPO training — that's the reinforcement learning method behind a lot of modern reasoning models — and they watched the power draw every half second.
Jane: Right, and they found something pretty wild. On a single GPU, they could train a little controller that cut power-limit violations by almost ninety percent while actually increasing the number of tokens the model produced. More output, less violation. That's not supposed to happen.
Tom: It sounds backwards, doesn't it? But the trick is that the controller isn't throttling the hardware. It's changing the training job itself — things like how many candidate answers the model generates per prompt. It spends the power budget where it buys output, and yields it where it doesn't.
Jane: And that's the key insight I want listeners to hold onto. Traditional power management is like putting a speed limiter on a car. This is more like teaching the driver when to press the gas and when to coast. Same car, same road, but you use less fuel.
Tom: I love that analogy, Jane. And the authors make the point that this matters because datacenters are hitting a wall. It's not silicon that's scarce anymore — it's electricity. The paper cites estimates that US datacenters are using over a hundred and seventy terawatt-hours a year, and that number is projected to double or triple.
Jane: So the question becomes, can we make these workloads smarter about power without slowing them down? And this paper says yes, at least at the single-GPU level. But then they try to scale it up to multi-GPU setups, and that's where things get really interesting.
Tom: Oh, we're definitely getting to that. Let's just say the title says "from one GPU to the fleet," and that journey is not a straight line. There's a twist in the middle that I think is going to surprise a lot of people.
Jane: It surprised me, that's for sure. Stick around, because we're about to dig into what happens when you try to control power on a model that's spread across four GPUs instead of one.
Summary: Tom: So we're back, and we're still on "Cutting AI Datacenter Energy with Reinforcement Learning." Jane, you teased the twist at the end of the last segment. Let's get into the meat of the paper now.
Jane: Alright, so the paper's structure is basically a progression. They start with a seven-billion-parameter model on one GPU. That works beautifully — the controller cuts violations by eighty-nine point eight percent, boosts token output by eighteen point one percent, and improves energy efficiency by twenty-six point two percent. Those are the headline numbers.
Tom: And then they go to a fourteen-billion-parameter model on two GPUs, and suddenly the controller does nothing. The violation rate barely moves. Then they try a seventy-two-billion-parameter model on four GPUs, and they replicate that null result three times with three different controller designs.
Jane: Right, and this is where I want to bring in Lu, because I think this is the most fascinating part of the paper. Lu, why does the controller just stop working when you shard the model?
Lu: Great question, Jane. The paper diagnoses it beautifully. When you shard a model across multiple GPUs, the GPUs execute as a pipeline. They take turns. So changing the group size — which is the knob that worked on a single GPU — changes how much work there is, but the pipeline just processes it at its own natural rate. The power draw doesn't move.
Tom: They measured it down to the tenth of a watt. Mean per-device power was one hundred forty-seven point zero watts with the controller and one hundred forty-seven point zero watts without it. Identical.
Lu: Exactly. And that's the "actuator authority" problem. The knob you're turning doesn't actually control anything anymore. It's like trying to steer a car by turning the radio dial. The radio changes, the car doesn't.
Jane: But here's the thing — they didn't give up. They did an actuator sweep, testing every software knob they could find on the 72B setup, and they found one that still worked. Generation concurrency — basically how many sequences you run through the pipeline at once — that moved power by seventeen to twenty-two percent.
Tom: So they rebuilt the controller around that knob, and this time it worked live on hardware. Across three replications, it produced thirty-five point seven percent more tokens than a static safe baseline, while keeping violations at around two point two seven percent against a one percent target.
Lu: And that's the honest part of the paper. They missed their own target. They set a one percent violation bar and landed at two point two seven percent. But the comparison that matters is against the uncontrolled baseline, which had a seventeen point seven two percent violation rate. They cut that by eighty-seven percent.
Jane: So the summary is: control works on a single GPU, fails on sharded models with the wrong actuator, and works again when you find the right actuator. It's a story about matching the control mechanism to the hardware reality.
Tom: And that's the kind of finding that only comes from actually running the experiments. You can't simulate your way to that conclusion. Which brings us to our next segment — what this means for actual datacenters and the grid.
Lu: And I think that's where the paper gets really provocative. Because once you know the workload is controllable, the question becomes how much capacity you actually need to build.
Improvements: Tom: We're back on "Cutting AI Datacenter Energy with Reinforcement Learning," and now we're getting to the part that I think has the biggest real-world impact. Jane, you mentioned the fleet analysis earlier. Let's bring in Meng for this one.
Jane: Absolutely. Meng, the paper does something clever at the end. They take the actual measured power traces from their experiments and compose them into a hypothetical sixteen-GPU fleet. Then they ask: how much capacity does this fleet actually need?
Meng: And the answer is about half of what the nameplate says. The fleet has a nameplate of six point four kilowatts, but the measured thirty-second peak demand is around three point one seven kilowatts with the best scheduling. That's a fifty percent provisioning fraction.
Tom: Which means you could oversubscribe the facility by roughly two times and never trip a breaker, as long as you stagger the job starts.
Meng: Right. And that's not a simulation. Those are real traces from real GPUs doing real training. The statistical multiplexing — the fact that jobs don't all peak at the same time — does most of the work. Scheduling and control add another six percentage points on top.
Jane: And this is where the economic numbers get big. The paper cites roughly four million dollars per megawatt for construction and interconnection. So for a one hundred-megawatt facility, going from one hundred percent provisioning to fifty percent provisioning is on the order of two hundred million dollars of avoided capital.
Tom: That's the kind of number that makes a CFO sit up straight. But Meng, you're the engineer. What's the catch?
Meng: The catch is that this is composed from traces, not measured on an operating fleet. The paper is honest about that. They call for a pilot — thirty to sixty days on a real mixed fleet, at an estimated cost of fifty thousand dollars — to verify the provisioning fraction.
Lu: And I'd add another catch. The fleet analysis assumes you can control the elastic jobs. The single-GPU jobs respond to the learned controller. The sharded jobs need the occupancy actuator. So the deployment architecture has to be right, or you don't get the savings.
Jane: But here's what I find genuinely exciting. The paper argues that AI training could become a dispatchable load — something the grid can ask to shed power during peak demand, the way they ask factories to shut down during heat waves. That's a huge deal for grid stability.
Meng: And it's not just about building less capacity. It's about using what you have more efficiently. The paper shows a twenty-six percent improvement in tokens per megawatt-hour on the single-GPU controller. That's more output for the same electricity.
Tom: So the improvements this paper suggests are really threefold: a controller that makes individual jobs power-aware, a scheduling strategy that exploits statistical diversity across jobs, and a provisioning framework that lets you build less and use what you build better.
Lu: And the carbon angle is real too. At the New York grid intensity, the efficiency gain translates to about twenty percent less carbon per million tokens. Scale that across the industry and you're talking about meaningful emissions reductions.
Jane: And that's the hook for our final segment — what this all means in the bigger picture, and whether the authors' vision holds up.
Conclusion: Tom: And we're wrapping up our discussion of "Cutting AI Datacenter Energy with Reinforcement Learning: Measured Power Control of LLM Training from One GPU to the Fleet." Jane, give us the final summary.
Jane: So the paper does three things. First, it characterizes the power behavior of GRPO training — the dominant post-training method for reasoning models — across three model scales. Second, it demonstrates a learned controller that works on a single GPU, fails on sharded models with the wrong actuator, and works again when you find the right one. Third, it composes the measured traces into a fleet analysis showing that mixed training fleets need roughly half their nameplate capacity.
Tom: And the headline numbers are worth repeating one more time. On the single GPU, eighty-nine point eight percent fewer violations, eighteen point one percent more tokens, twenty-six point two percent better energy efficiency. On the live 72B rollout workload, thirty-five point seven percent more output than the static safe baseline and eighty-seven percent fewer violations than uncontrolled operation.
Lu: And the deeper lesson, I think, is about matching control to reality. The paper shows that a knob that works at one scale can be useless at another. You have to measure, not assume. That's the scientific discipline that makes the results trustworthy.
Meng: From an engineering standpoint, the pilot they propose is the right next step. Fifty thousand dollars to verify a two-hundred-million-dollar savings is a no-brainer. I'd sign off on that.
Jane: And I think the cultural impact is worth noting too. This paper reframes AI training from an inflexible load that the grid has to accommodate to a flexible resource that can help the grid. That's a shift in how we think about the relationship between AI and energy.
Tom: Lalam, you've been quiet. What's your take?
Lalam: I think the most impactful vision here is the one where AI infrastructure becomes a partner in grid management rather than a burden on it. If training workloads can shed load on demand without losing output, then datacenters can participate in demand-response programs, which means the grid can absorb more renewable energy without building as much new capacity. That's a win for the climate and for the economics of AI.
Tom: That's a beautiful way to put it. So we've got measured results, a clear failure mode, a fix, and a path to validation. That's a complete paper.
Jane: It really is. And it's the kind of work that makes you optimistic about the future of AI infrastructure. We're not just building bigger — we're building smarter.
Tom: And on that note, we're saying goodbye to this paper. Thanks for joining us, and we'll see you next time with something new from the arXiv.
Jane: Take care, everyone.
Eliseo Curcio
Advanced Department of Artificial Intelligence and Energy – New York
cs.AI, cs.SY, eess.SY
Submitted: 2026-07-27
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 58/100
The gist: This paper presents the first power characterization of reinforcement-learning post-training (specifically GRPO training) and demonstrates a learned power controller that adapts the workload's own
Key concepts
- Reinforcement Learning (RL)
- A method used to train a controller that learns optimal power management strategies for LLM training. Instead of merely simulating, the system adjusts the training job itself—like candidate answers per prompt—to spend power where it maximizes output.
- Sharding/Single GPU vs. Fleet
- The paper tested power control on both a single GPU and across multiple GPUs (a fleet). The hosts noted that the controller worked well on one GPU but failed when the model was spread across multiple GPUs, requiring a different control mechanism.
- Generation Concurrency
- A specific software knob found by the researchers that allowed power control to work on sharded models. It refers to how many sequences are run through the pipeline at once, which successfully moved power draw and improved efficiency.
- Provisioning Fraction
- The ratio of actual measured peak demand to the facility's nameplate capacity. The paper suggests that due to statistical multiplexing, AI fleets could operate with a much lower provisioning fraction (e.g., 50%), saving billions in construction costs.
Terminology
Summary
This paper presents the first power characterization of reinforcement-learning post-training (specifically GRPO training) and demonstrates a learned power controller that adapts the workload's own generation parameters to measured power. The study instruments GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s, comprising approximately 200,000 GPU telemetry samples at half-second resolution in the original BF16 campaign and over 380,000 samples in total once the generation-proxy campaign is included.
Single-accelerator results (7B on one A100): The 500-step baseline run shows generation draws a sustained 250 to 315 W with a mean of 299.2 W and brief peaks to 410.6 W, against an idle draw of 86 W. Against the evaluation cap of 314.6 W, the baseline spends 43.2% of samples in violation. A PPO meta-controller trained on this trace reduces power-limit violations by 89.8% (from 216 to 22), increases token output by 18.1% (from 1,731,712 to 2,044,928 tokens), reduces mean power by 19.3 W and peak power by 35.8 W, and improves energy efficiency by 26.2% in tokens per megawatt-hour. Energy intensity falls from 1.026 to 0.813 Wh per thousand tokens, a 20.8% improvement.
Multi-GPU null results (14B and 72B): At 14B on two A100s, the controller arm records 8.10% violation rate against a 7.99% baseline, indistinguishable from zero.
At 72B on four A100s, three different controllers (transferred multi-scale policy, cluster-trained variant, and signal-safe design) change the violation rate by +4.4%, −1.6%, and −2.7% respectively, three essentially null results.
The null is verified as an integration failure of the actuator, not absent actuation: the controller's group size moved through sequences such as five, six, five, four, three, two in response to measured power. The agent acted; the power did not move.
Mean per-device power is 147.0 W with and without the controller, identical to the tenth of a watt.
Mechanism of the scale boundary: Three measured facts explain the boundary. First, "the actuator loses authority under sharding: mean per-device power is 147.0 W with the controller and 147.0 W without it... because a pipeline of sharded layers executes at its natural rate regardless of how many candidate answers each prompt receives. Second, the disturbance changes character: the pairwise correlation of the four devices' power draws is −0.13, and cluster peaks are
brief coincidences of pipeline stages, with a median duration of 0.5 seconds and a ninetieth percentile of 2 seconds. Third, the control loop is mismatched:
the controller acts once per training step, approximately every 55 seconds... while the disturbances it would need to reject live below two seconds."
Actuator-authority sweep: The sweep of every user-space knob on the 72B condition shows that generation batch size (concurrency) moves mean cluster power by 22.5% and the 30-second peak by 24.0% at 5.2x throughput; group size applied as concurrency moves mean power by 16.9% and the peak by 17.0% at 3.3x throughput; completion length moves power by only 3.5%; inter-batch pause by 6.3%; and firmware power caps/clocks were denied by the cloud container with insufficient permissions.
The principle: a knob has power authority precisely insofar as it changes pipeline occupancy rather than work volume.
Rebuilt occupancy controller: A controller was rebuilt around the generation batch level as the actuator, trained with constrained reinforcement learning (Lagrangian dual ascent against a 1% violation-rate target) on measured transition traces from a randomized calibration run. The simulation study found a phase boundary: at a 500 W budget, "no tested PPO policy or cyclic mixing schedule satisfied the pre-specified criterion of at least 10% more tokens than the static batch-4 baseline at 1% violations or fewer... a controller that correctly identifies inaction as optimal validates the method." At a 520 W budget, the selected controller achieved 94,349 tokens at 0.49% violations, a 40% gain over static batch-4.
Live four-arm hardware study: Four arms ran live on four A100s at equal twenty-minute durations with three replications each across two physical pods, at a 561 W / 30-second budget. The PPO occupancy controller produced 183,069 ± 9,093 tokens (154.8 ± 7.3 tokens/s) at 2.27% ± 1.08% violation rate, versus static batch 8 at 222,472 ± 5,549 tokens at 17.72% ± 11.40% violations, static batch 4 at 134,860 ± 10,607 tokens at 0.01% ± 0.02% violations, and optimized hysteresis at 166,189 ± 9,106 tokens at 1.96% ± 0.57% violations. The controller produced 35.7% more tokens than the static safe baseline
and 87.2% fewer violations than uncontrolled operation,
with the highest mean throughput and the lowest mean energy per token among the constrained controllers.
The controller missed the strict 1% violation bar (2.27%), and the heuristic had a lower mean violation rate and won the third replication.
Measurement-window analysis: Recomputing the 72B record as a rolling mean over increasing windows collapses the violation rate from 23.6% at the instantaneous half-second resolution, through 10.8% at ten seconds and 1.6% at thirty seconds, to exactly zero at five minutes, with no controller involved at all.
The composed fleet falls from 9.1% instantaneous to zero at every window of thirty seconds and longer.
The analysis divides violations into plateau violations
(sustained draw above the limit, which survive any averaging and are eliminated by the controller) and transient violations
(sub-second pipeline overlaps, mostly averaged away at the integration times infrastructure actually uses
).
Fleet-level provisioning analysis: A representative fleet of sixteen A100s (two 72B jobs on four accelerators each, two 14B jobs on two each, and four 7B jobs on one each, 6.4 kW aggregate nameplate) composed from real measured traces shows: naive aligned starts give a 30-second peak of 3.59 kW (56% of nameplate); random starts 3.44 kW (54%); phase-staggered 3.38 kW (53%); and staggered with job-level control 3.17 kW (50%). Across two independent trace compositions the final figure ranged from 45 to 50% of nameplate.
The claim: for this fleet mix, measured behavior supports approximately twofold oversubscription of nameplate capacity with zero violations at any measurement window of thirty seconds or longer.
Economic and carbon implications: At the conservative 50% provisioning fraction, "a facility designed around 100 MW of accelerator nameplate requires roughly 50 MW of provisioned capacity, which corresponds to on the order of two hundred million dollars of avoided capital and several years of queue position per 100 MW." The efficiency channel, computed as capacity times 70% utilization times 8,760 hours times a 50% controllable fraction times the 20.8% fixed-output saving, yields approximately 63,800 MWh, 5.1 million, and 15,900 tonnes of CO2 per year for a 100 MW facility. Applied to New York State's multi-gigawatt interconnection pipeline, the capacity channel implies on the order of 2,300 MW of avoidable new construction and roughly 9.2 billion dollars of capital,
and the efficiency channel corresponds to approximately 2.9 TWh, 235 million dollars, and 733,000 tonnes of CO2 avoided annually... equivalent to removing roughly 159,000 passenger cars from the road.
Signal-protection study: Reducing group size erodes the statistical signal GRPO learns from: the fraction of groups carrying zero reward variance reaching 0.94
at group size two. Two reward-based protections failed (reward bonus for signal quality caused the policy to maximize group size and ignore power; reward penalty on signal collapse produced the mirrored failure), while a structural action-space floor at group size three achieved a 65.3% violation reduction at a 4.2% token gain. The paper notes this converges independently with Meta's Autodata design: both systems arrived independently at the same design conclusion, that the learning signal must be protected by a hard constraint and not by a reward term.
Controller design evolution: Four controller generations are reported. G1 carried raw power values and failed to transfer. G2 replaced raw watts with cap-relative quantities and sampled episodes jointly from all three trace environments, producing a single policy that transfers across scales without retraining.
G3 replaced the fixed violation penalty with a Lagrangian formulation with dual ascent, requiring feasibility measurement before target selection (a 10% target lay outside the feasible band of the 14B environment and drove the multiplier to divergence; the 8% target inside the band produced textbook behavior). G4 added the structural group-size floor.
Key methodological findings: Trace-replay simulators with parametric actuator models systematically overestimate control authority, and hardware validation is mandatory.
The replay environment credited controllers with up to 65% violation reduction at 72B scale while hardware showed approximately zero. The paper also specifies a three-layer deployment architecture: job-level reinforcement-learning control (acting once per training step for elastic single-device jobs, at generation-batch boundaries for sharded jobs), firmware power caps in milliseconds for transient events, and a fleet scheduler over minutes for job-level staggering.
Limitations: The occupancy-controller results use an AWQ-quantized generation proxy rather than the BF16 GRPO training loop; the live controller missed the 1% violation bar; one of three training seeds passed validation; the headline 7B controller result is trace evaluation rather than live in-loop hardware; the 50% provisioning fraction is composed from real traces rather than measured on an operated fleet; and the economic conversions rest on published figures. A validation pilot is specified: "an instrumented pilot on one operating mixed fleet... for thirty to sixty days at an estimated cost of fifty thousand dollars, with a pre-specified success criterion of at least a 30% reduction in required provisioning relative to nameplate at zero violations on any thirty-second window of the operator's own metering."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:
Improvement: Integrate a PPO-based meta-controller that observes normalized power telemetry (cap-relative, not raw watts) and dynamically adjusts GRPO training parameters (group size, batch size) at each training step.
What the improved system can do:
-
Reduce power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh) on single-GPU 7B training
-
Transfer zero-shot across model scales (7B, 14B, 72B) without retraining, because state features are normalized to the power cap
-
Maintain training signal quality via a hard action-space floor (group size ≥ 3), preventing the 94% zero-reward-variance collapse observed at group size 2
The improved system is a power-aware, self-adapting training infrastructure that:
-
Learns to control its own power consumption via RL, without hand-tuned penalties
-
Transfers across model scales and hardware configurations via cap-normalized state representations
-
Protects training signal quality via structural constraints, not fragile reward shaping
-
Distinguishes actionable violations from harmless transients via measurement-window analysis
-
Provisions fleets at 50% of nameplate, saving 200M per 100 MW of avoided capacity
-
Reduces energy per token by 21% and increases tokens per MWh by 26%
-
Validates all claims on real hardware telemetry, with disclosed limitations and a concrete pilot path
This makes AI training a dispatchable, grid-flexible load rather than an inflexible consumer, enabling demand response, oversubscription, and significant capital and carbon savings.
Abstract
Reinforcement-learning post-training dominates modern language-model development, yet its power behavior on GPU hardware has not been characterized, and datacenters manage GPU power with workload-blind mechanisms, static caps and reactive throttling, that slow hardware indiscriminately. We instrument GRPO training with half-second power telemetry at 7B, 14B, and 72B scales on one to four A100s (380,000+ samples), and train a PPO meta-controller that adapts the workload's own generation parameters to measured power. Against the full 500-step 7B trace, the controller cuts power-limit violations by 89.8% while increasing token output by 18.1% and energy efficiency by 26.2% (tokens per MWh). Deployed live at 72B, the same controller family yields replicated null results, diagnosed as the group-size actuator losing authority under model sharding. An actuator-authority sweep shows the same parameters applied as generation concurrency retain 17-22% power authority, isolating an occupancy-versus-volume principle; a controller rebuilt on that actuator controls a live 72B rollout-generation workload across three replications: 35.7% more output than a static safe baseline at 2.27 +/- 1.08% budget violations, 87.2% fewer violations than uncontrolled operation, and the best mean throughput and energy per token among constrained controllers, with an adaptive threshold rule matching it in one of three operating conditions. Under realistic measurement windows the original 72B transients fall from 23.6% at half-second resolution to 1.6% at 30 s and zero at 5 min; a composed 16-GPU fleet shows zero violations at 30 s and longer, with peak demand at 50-56% of nameplate. For this fleet mix, roughly twofold oversubscription of nameplate appears feasible, subject to operator validation. We quantify the economic and carbon consequences and specify a low-cost operator pilot.
Sources
- Power Stabilization for AI Training Datacenters
- Secrets of RLHF in Large Language Models Part I: PPO
- Gymnasium: A Standard Interface for Reinforcement Learning Environments
- Autodata: An agentic data scientist to create high quality synthetic data
- GPT-4 Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Qwen2.5 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection