The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution".
Jane: The paper was written by Deepak Panigrahy and Aakash Tyagi from Independent Researcher and Texas A&M University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1: Tom: Welcome back to the show, everyone! Today we're digging into a paper that's got a really punchy title: "The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution." Jane, I have to say, when I first read that title I thought, wait, NVIDIA's flagship edge hardware can't measure its own energy? That seems like a pretty big deal.
Jane: It is a big deal, Tom, and honestly the title undersells it a bit. This paper is by Deepak Panigrahy and Aakash Tyagi, and they're basically saying that if you buy NVIDIA's new GB10-based systems — the DGX Spark, the ASUS Ascent GXten all those desktop AI boxes that shipped in two thousand twenty-six — you get zero ability to measure how much energy the CPU is using. Zero. No counters, no sensors, nothing through any supported software interface.
Tom: And that matters because of their earlier work, right? They published this framework called A-LEMS where they showed that agentic AI workflows — you know, when an AI has to plan, call tools, retry failed steps — those consume over four times more energy per successful goal than simple linear AI tasks. The orchestration overhead, the thinking between steps, that's where the energy goes.
Jane: Exactly. And to measure that, you need to attribute energy to specific processes. On Intel xeighty-six chips, there's this thing called RAPL — Running Average Power Limit — that's been giving researchers hardware energy counters since two thousand eleven. You can literally read how many joules a process consumed. But on ARM chips like the GB10, there's no equivalent. And this paper systematically proves that.
Tom: They went through every possible interface, didn't they? I mean, they checked SCMI, they checked I2C buses, they checked IPMI, they checked hwmon, they checked power supply sysfs. Seven different interfaces, and the only energy telemetry on the whole platform is GPU power through NVML.
Jane: Right, and here's the kicker. The GPU power reading is instantaneous power, not cumulative energy. So you can see how many watts the GPU is drawing right now, but you can't sum up energy over time. And the CPU — where all the orchestration logic runs — has nothing at all. The paper calls it "energy-dark." The platform has rich performance counters, thermal zones, DVFS controls, but zero energy counters for the CPU.
Tom: So if you're a researcher trying to replicate their OOI measurements — the Orchestration Overhead Index — on this hardware, you just can't do it. The paper's earlier work showed orchestration overhead ranging from zero point six two times to twelve point six eight times depending on the task. But you can't reproduce that on the GB10.
Jane: And that's the "blind spot" in the title. It's not that the hardware can't measure energy — we'll get into that later — it's that NVIDIA chose not to expose it. The data exists inside the chip, but there's no supported way to read it. For researchers studying energy-efficient AI, that's a wall.
Tom: I love that they actually checked whether the firmware already computes this data internally. And spoiler alert: it does. But we'll save that for the next segment. Jane, what do you think is the most frustrating part for the research community here?
Jane: Honestly, Tom, it's the asymmetry. On xeighty-six you have RAPL and you can do per-process attribution. On the GB10, you have nothing. So any energy research done on this hardware has to rely on external power meters and statistical estimation, which introduces uncertainty. The paper makes the point really clearly: the measurement gap is structural, not absolute.
Tom: And with seven OEMs shipping GB10 systems — Dell, HP, ASUS, MSI, Acer, Gigabyte — that's a lot of energy-blind hardware going out into the world. Stick around, because next we're going to talk about the smoking gun they found — the firmware already has the energy data, it's just locked away.
Paper discussion segment 2: Tom: Welcome back! We're talking about "The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution." Jane, last segment we established the problem — no CPU energy counters on the GB10. But this paper has a really interesting twist: the data actually exists inside the chip. The firmware is already computing it.
Jane: That's the part that makes this paper feel almost like a detective story. They found that the MediaTek firmware on the GB10 runs something called a System Power Budget Manager — SPBM for short. It's a shared memory region that gets updated every one hundred milliseconds or so, and it contains per-rail power readings and cumulative energy accumulators for the P-core cluster, the E-core cluster, the GPU, and the whole SoC package.
Tom: So the chip knows exactly how much energy each domain is using. It's just not exposed through any supported interface. They discovered this because a community developer reverse-engineered an undocumented ACPI interface to read that data. And when they confirmed it on a second machine — the Acer Veriton GN100 — the energy accumulators were live and incrementing.
Jane: Right, and here's the thing that really gets me. They also checked the SCMI bus — that's ARM's standard interface for system management. The bus is active on the GB10. Four protocol drivers are loaded. But the powercap and sensor protocols are missing. The kernel support for reading SCMI energy data is being actively developed by Linux maintainers, but the firmware just doesn't expose it.
Tom: And NVIDIA's official response? The paper quotes them saying there are "no plans to expose CPU rail information." So the measurement infrastructure exists, the kernel support is being built, but the vendor decision is to keep it locked.
Jane: Exactly. And that's why the paper argues this is a product prioritization decision, not a silicon limitation. The PMIC — the power management IC — has to monitor per-rail current for DVFS and thermal protection. That's how the chip decides to boost or throttle cores. So the current sensing is already there. The energy integration is already there. The only missing piece is the firmware exposing it.
Tom: They make a great comparison to NVIDIA's own Jetson Orin platform, which has three INA3221 power monitors on the board giving per-rail power to userspace. So NVIDIA has done this before on ARM hardware. The GB10 is actually a step backward in energy observability compared to their cheaper platform.
Jane: And that's the frustrating part for researchers. The paper's earlier work on agentic AI energy — the four point three three times orchestration overhead finding — depends on per-process energy attribution. On the GB10, you can't do that for the CPU. You can measure GPU power, but the orchestration logic runs on the CPU. So the most important energy cost in agentic workloads is invisible.
Tom: They also cite another paper — Raj et al. — showing that CPU-side processing accounts for up to ninety point six percent of total latency and forty-four percent of total dynamic energy in agentic workloads. So the CPU is where the energy goes, and the CPU is exactly what you can't measure.
Jane: Right. And there's a regulatory angle too. The EU AI Act requires GPAI model providers to document energy consumption. If you're deploying agentic models on GB10 hardware, you have no supported way to measure that. The paper argues this creates a compliance gap, even if it's not technically a violation.
Tom: So we've got a chip that knows its own energy consumption but won't tell anyone. What do the authors propose? I know they have a hardware requirements spec and a calibration bridge. That's coming up next.
Jane: And I'll tease it this way — they think the fix could be as simple as a firmware update. No hardware redesign needed. The data is already there.
Paper discussion segment 3: Tom: We're back with "The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution." Jane, the authors don't just complain — they actually propose solutions. Let's talk about what they think should happen.
Jane: Right, and they break it into three layers. First, they propose a hardware requirements specification. Any platform that wants to support process-level energy attribution for agentic AI needs per-power-domain cumulative energy counters, readable from userspace with low latency and millijoule resolution. You need enough granularity to separate CPU, GPU, DRAM, and I/O. And you need monotonically increasing counters with defined overflow semantics.
Tom: That's the ideal. But what do you do right now, before NVIDIA ships a firmware update?
Jane: They propose what they call an interim calibration bridge. You use the native GPU power reading from NVML, you add an external DC power meter inline at the power jack, and then you derive the CPU-plus-system energy by subtracting GPU power from total power. It's a two-channel decomposition — GPU versus everything else.
Tom: And they actually validated that on the Acer Veriton GN100, right? They got the SPBM interface working there and confirmed the energy accumulators were live.
Jane: Yes, and this is where it gets really interesting. On the GN100, the SPBM interface gives you total wall power as a native sysfs channel. So you don't even need the external meter. And they cross-validated the GPU energy using DCGM field one hundred fifty-six which is NVIDIA's official cumulative energy counter. At idle, they saw about four thousand four hundred thirty-seven millijoules per second at four point four watts, which is consistent with the instantaneous power reading.
Tom: So the data is all there on the GN100. The GXten had some kind of memory conflict that blocked the community driver, but the GN100 worked cleanly. That's a pretty strong proof that this isn't a silicon limitation.
Jane: Exactly. And then the third layer is the standards-track solution. They want NVIDIA to expose SCMI Powercap domains for CPU cluster, GPU, DRAM, and SoC total, with cumulative energy readout in microjoules. The kernel support is being built — there are patches from February and April two thousand twenty-six adding SCMI powercap measurement averaging interval support. What's missing is the firmware decision to enable it.
Tom: And the cost to fix this? The paper says zero. No board redesign, no additional components. The PMIC already has the current sensors. The firmware already computes the energy. It's literally a firmware update to expose what's already being calculated internally.
Jane: Right. And they even note that a hardware solution — adding an INA3221 like the Jetson Orin uses — would cost about.50 per board. So even the expensive fix is cheap. The issue is purely a product decision.
Tom: They also have a nice grading framework in the paper. The GB10 gets a "LIMITED" grade. With the community driver on the GN100, it's "PARTIAL." With SCMI powercap enabled, it would be "MEASURED" — the same grade as Intel xeighty-six with RAPL.
Jane: And that's the goal. They want the GB10 to be the first ARM desktop AI platform with research-grade energy attribution. Given that NVIDIA is explicitly marketing the DGX Spark to researchers and developers, that would be a competitive advantage.
Tom: So the message to NVIDIA is pretty clear: the measurement infrastructure exists, the kernel support is coming, and the only thing standing between the research community and reproducible energy measurement is a firmware decision. I hope they're listening.
Jane: Me too, Tom. And there's one more angle we haven't touched — what this means for the broader ecosystem. That's where we're heading in the final segment.
Conclusion: Tom: We're wrapping up our discussion of "The Energy Blind Spot: NVIDIA's Flagship Edge AI Hardware Cannot Support Process-Level Energy Attribution." Jane, let's pull it all together for our listeners.
Jane: So the paper makes three core points. First, agentic AI workloads — where an AI plans, calls tools, retries — consume dramatically more energy than simple inference, and the orchestration overhead happens mostly on the CPU. Second, NVIDIA's flagship edge AI hardware, the GB10, provides no supported way to measure CPU energy. Third, the measurement infrastructure actually exists inside the chip — the firmware computes per-rail energy every one hundred milliseconds — but NVIDIA has decided not to expose it.
Tom: And that decision has real consequences. Researchers can't replicate energy measurements across platforms. Developers can't optimize their agentic systems because they can't see where the energy goes. Organizations facing regulatory reporting requirements can't document energy consumption on this hardware.
Jane: The paper's constructive message is that this is fixable. A firmware update exposing SCMI powercap would move the GB10 from "LIMITED" to "MEASURED" in their grading framework. The kernel support is being built. The hardware already has the sensors. It's a vendor decision, not a technical constraint.
Tom: And they're calling on the low-carbon computing community to demand energy observability as a first-class hardware requirement. Not an afterthought, not a nice-to-have, but a fundamental capability that any AI hardware should provide.
Jane: Right. Because if we're serious about making AI more energy-efficient, we need to be able to measure it. You can't optimize what you can't see. And right now, on the most popular edge AI hardware shipping in two thousand twenty-six the energy consumption of the CPU — where all the orchestration logic runs — is invisible.
Tom: Well said, Jane. That's a great note to end on. Thanks to everyone who listened. We'll be back next episode with another paper from the arXiv. Until then, keep measuring what matters.
Jane: Goodbye, everyone!
Deepak Panigrahy, Aakash Tyagi
Independent Researcher · Texas A&M University
cs.LG, cs.AI, cs.AR, cs.DC, cs.PF
Submitted: 2026-08-14
Updated: 2026-08-18
Code: https://github.com/antheas/spark_hwmon
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 92/100
Key concepts
- Agentic AI Workflows
- These are complex AI tasks where the system must plan, call external tools, and handle failures by retrying steps. The energy consumed during this planning and orchestration phase is significantly higher than simple linear tasks.
- Process-Level Energy Attribution
- This is the capability to measure precisely how much electrical energy a specific software process uses. On certain hardware, like Intel's x86 chips using RAPL, this is possible; on the GB10, it is not supported by any interface.
- SPBM (System Power Budget Manager)
- This is an internal mechanism within the MediaTek firmware on the GB10 chip. It tracks per-rail power readings and cumulative energy accumulators for various components, like the CPU clusters and GPU. The data exists here but is not exposed to userspace.
Terminology
Summary
Summary
This paper reports a systematic energy-observability audit of the ASUS Ascent GX10, a GB10-based edge AI system, and finds that the platform exposes no CPU energy counter, no INA power-rail monitor, no IPMI/BMC, and no SCMI powercap protocol through any supported software interface. The only on-device energy telemetry is instantaneous GPU power via NVML.
The audit was performed on the ASUS Ascent GX10 (firmware GX10DGX.0104.2026.0326.1657, kernel 6.17.0-1018-nvidia, Ubuntu 24.04.4 LTS, aarch64), containing a MediaTek-designed SoC with 10× ARM Cortex-X925 performance cores and 10× ARM Cortex-A725 efficiency cores coupled via NVLink-C2C to an NVIDIA Blackwell GPU with 128 GB unified LPDDR5X memory. Seven energy measurement interfaces were probed: ARM SCMI powercap (bus active, but only clock, regulator, and MPAM drivers loaded; no powercap or sensor), ARM PMU energy events (zero hardware energy events), INA3221/INA226 (zero devices on all six I2C buses), IPMI/BMC (no /dev/ipmi0), hwmon energy/power (no results; only temperature sensors), power supply subsystem (empty), and NVML (GPU power only, 3.84 W avg, no cumulative energy counter).
The paper's central finding is the SCMI smoking gun
: the SCMI protocol bus is registered and active, with four protocol drivers loaded (scmi-clocks, scmi-regulator, scmi-mpam-driver, scmi-imx-bbm-key), but neither scmi-powercap nor scmi-sensor drivers are present. The authors argue this is a firmware decision, not a silicon limitation. They discovered that the MediaTek SSPM firmware maintains a System Power Budget Manager (SPBM) shared memory region, populated at approximately 100 ms intervals, containing per-rail power in milliwatts, cumulative energy accumulators in millijoules (covering P-core cluster, E-core cluster, GPU, and total SoC package), and per-zone temperatures, accessible via an undocumented ACPI DSM method on device NVDA8800. NVIDIA has officially stated there are no plans to expose CPU rail information.
The paper establishes that the GB10 could expose energy data through three layers of evidence: (1) the PMIC must monitor per-rail current for DVFS decisions, so the data exists; (2) the SCMI bus is ready, with kernel support for SCMI powercap MAI being actively developed; (3) NVIDIA's own Jetson Orin platform includes INA3221 power monitors providing per-rail energy, making the GB10 strictly less capable despite being more expensive. The cost to fix is zero—a firmware update exposing existing current-sense data through SCMI powercap would suffice.
The paper contrasts the GB10's rich telemetry (70+ PMU events, 7 thermal zones, full DVFS control, GPU utilization) with its complete absence of energy counters. A 5-second idle sample shows the X925 cluster at 332M cycles and the A725 cluster at 174M cycles, but no energy values. The authors note that performance counters are useless for energy attribution
without ground-truth energy measurement to calibrate against.
The paper discusses attribution without on-device counters, citing FaasMeter's Shapley-value-based disaggregation approach, but identifies limitations: unbounded error without calibration, multi-process attribution approximation, reproducibility concerns, and overhead from external DC metering. The authors conclude that attribution is possible but with bounded uncertainty, coarser resolution, and higher operational overhead than hardware counter-based methods.
The paper proposes a hardware requirements specification: per-power-domain cumulative energy counters readable from userspace at < 1 ms read latency with resolution ≤ 1 mJ, sufficient domain granularity to separate CPU, GPU, DRAM, and I/O, and monotonically increasing counters with defined overflow semantics. It grades platforms: Intel x86 + RAPL (MEASURED), NVIDIA Jetson Orin (MEASURED), NVIDIA GB10 (LIMITED), GB10 + spark hwmon on Acer GN100 (PARTIAL), Apple M-series (LIMITED), Qualcomm Snapdragon (LIMITED).
An interim calibration bridge is proposed: native NVML GPU power, external DC metering at ≥ 1 kHz sampling, and derived attribution via E cpu+sys = E total − E gpu. Post-submission validation on the Acer Veriton GN100 confirmed SPBM binding with 45/45 DSM offsets resolved, 14 power channels, and 4 cumulative energy accumulators (cpu p, cpu e, pkg, gpu) incrementing, with idle rates of cpu p ≈ 617 mW, cpu e ≈ 156 mW, gpu ≈ 5,354 mW. DCGM field 156 (total energy consumption) was also validated, providing official cumulative GPU energy.
The standards-track solution is SCMI powercap: the GB10's SCP firmware should expose SCMI Powercap domains for CPU cluster, GPU, DRAM, and SoC-total, with cumulative energy readout in microjoules and configurable measurement averaging intervals. The kernel infrastructure is being built (Radford's MAI patches, SCMIv4.0 powercap extensions); what is missing is the firmware-side decision.
The paper frames this as a systemic ARM ecosystem gap, not product-specific: Apple M-series, Qualcomm Snapdragon, and Ampere Altra all lack userspace-accessible per-domain energy counters. The edge paradox
is that edge AI is motivated by energy efficiency, but the hardware enabling it removes the measurement infrastructure needed to verify savings. Regulatory exposure is discussed: the EU AI Act (Regulation 2024/1689) requires GPAI model providers to document energy consumption under Annex XI, Section 1(e), and California's SB 253 requires Scope 1–3 emissions disclosure starting 2026, creating compliance environments where accurate, reproducible energy measurement is needed.
The paper's limitations include: audit coverage of only two GB10-based systems (ASUS Ascent GX10 and Acer Veriton GN100), board-specific I2C scan results, firmware-version-specific SCMI findings, and the unsupported nature of the community spark hwmon driver. The authors encourage replication on DGX Spark, Dell ProMax, and HP ZGX Nano.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Add an energy-attribution layer to agentic AI orchestration frameworks that detects when hardware lacks energy counters and automatically switches to estimation mode.
What the improved system can do:
-
Detect absence of RAPL/SCMI powercap interfaces at runtime
-
Automatically fall back to the paper's calibration bridge:
E cpu+sys = E total - E gpuusing external DC metering or SPBM (where available) -
Track per-process CPU time via
/proc/ pid /statand scale against the derived CPU+system energy -
Flag measurement confidence levels (MEASURED vs. LIMITED vs. PARTIAL) in all energy reports
-
Prevent silent TDP-based fallbacks that introduce unbounded error
These improvements collectively enable: reproducible energy accounting on ARM edge hardware, regulatory compliance documentation, orchestration overhead detection and optimization, and honest cross-platform energy comparisons — all without requiring hardware changes, only firmware-level cooperation and software adaptation.
Sources
- Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
- Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
- FaasMeter: Energy-First Serverless Computing
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks