ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs

arXiv:2610.00267 · eess.SY, cs.SY · Submitted 2026-09-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs".

Dev: Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path.

Rosa: First, who's behind it and why it matters.

Paper summary: Rosa: To summarize what we just covered, this paper introduces ThermE as a runtime system specifically designed to predict and manage the shared thermal headroom for sustained LLM inference when the CPU, GPU, and RAM all share a single cooling path.

Dev: The central idea is that existing vendor governors only react when thermal limits are almost hit, whereas ThermE attempts to provide a control decision that improves performance but respects future thermal constraints by predicting what happens next.

Taro: So the thesis really boils down to creating a runtime system that moves from reactive throttling based on current heat levels to proactive management based on forecasted demand and coupled thermal dynamics.

Rosa: Exactly, and it claims this is achieved through its Fast LLM-to-Heat Compiler, which maps the model and requests to domain heats without executing the LLMs first.

Dev: Then there's the PDE-Constrained Headroom Predictor that leverages ThermPINN for offline thermal identification of those coupled dynamics, and online, it uses a Reduced Headroom Predictor to give uncertainty-calibrated estimates of what's left.

Taro: I see the value in using PINNs here because it lets them model those complex physical interactions—the cross-domain heat propagation that happens when one part gets hot and affects the others.

Rosa: Right, and on top of that, they have an Uncertainty-Aware Action Scheduler that uses a beam search to pick actions based on a multi-objective function focused on serving quality and maintaining a long runway in normal operation.

Dev: It’s designed to balance achieving SLO-compliant completed tokens while actively penalizing the cumulative shared-headroom deficit over time, which is how it manages the resource.

Taro: That focus on balancing immediate service progress against the long-term thermal budget sounds like a very practical approach for autonomous systems that have finite operational lifespans in remote locations.

Rosa: It’s about making sure that every token generated contributes positively to both performance and future thermal stability, which is what the authors claim is the core contribution of ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs.

Conclusion: Rosa: Thinking about the title, "ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs," it really captures the essence of what this research is attempting to do.

Dev: I think the authors are pointing toward a future where we have runtime systems that can handle thermal constraints not just by reacting, but by intelligently planning ahead based on predicted workloads.

Taro: The implication here for autonomous agents is that we might see AI deployed in satellites or remote devices running much longer because they aren't constantly fighting against immediate thermal limits.

Rosa: Right, and it suggests that the impact could be significant in making continuous, reliable intelligence possible on these compact platforms where cooling is naturally limited.

Dev: It moves the discussion from just hardware limitations to a software layer that can manage those constraints proactively throughout the inference process.

Taro: If this system proves robust across different hardware profiles, it could allow us to design AI deployments with much more realistic operational envelopes in mind.

Rosa: So essentially, ThermE is providing a framework for managing shared thermal resources to ensure serving quality is maintained over the long term, which is what the paper about "ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs" accomplishes.

The Hong Kong Polytechnic University · Southern University of Science and Technology

eess.SY, cs.SY

Submitted: 2026-09-24

Updated: 2026-09-24

Comments: 13 pages, 16 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path.

Key concepts

Fast LLM-to-Heat Compiler
This component analyzes different types of LLM tasks (like MMA or CUDA) and maps them to specific demands on compute and memory pressure. It does this without running the actual LLM, creating a structured representation of the workload's thermal requirements for the next step.
PDE-Constrained Headroom Predictor
This uses physics-informed neural networks (ThermPINN) to model how heat moves through the shared package over time. It predicts future thermal headroom by solving a complex heat transfer equation, providing uncertainty-calibrated estimates of available cooling capacity.
Uncertainty-Aware Action Scheduler
This component selects the best sequence of operations by using a beam search. It chooses actions that maximize serving goals (like completed tokens) while penalizing future thermal deficits, ensuring the system plans for sustained performance rather than just immediate output.

Terminology

Summary

Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. ThermE is a runtime system that predicts and jointly manages shared thermal headroom for sustained LLM inference on edge SoCs.

How it works

ThermE consists of three coupled components designed to proactively manage shared thermal headroom:

  1. Fast LLM-to-Heat Compiler: This component maps the model, requests, runtime state, and candidate actions to domain heats without executing LLMs. It extracts four types of static workloads—MMA, CUDA, CPU servingruntime workloads—and uses a Resource Pressure Encoder to map these demands to compute and memory pressure vectors.

  2. PDE-Constrained Headroom Predictor: This component uses ThermPINN for offline thermal identification of coupled thermal dynamics governed by a transient heat-transfer PDE (Eq. 5). Online, the Reduced Headroom Predictor (RHP) uses these operators to produce uncertainty-calibrated headroom estimates with low computational overhead.

  3. Uncertainty-Aware Action Scheduler: This component selects actions that balance serving quality and future headroom. It uses a beam search to select the highest-scoring feasible sequence based on a multi-objective function, maximizing terms like SLO-compliant completed tokens and rewarding longer runway in normal operation while penalizing the cumulative shared-headroom deficit.

Key Insights from Motivation Experiments

The paper characterizes sustained LLM inference under thermal constraints through three motivational experiments that revealed critical properties:

  1. Built-in thermal throttling is reactive and preset, activating only after most thermal headroom has been consumed, which degrades serving performance through staged frequency reductions. This motivates control decisions that account for future headroom before hardware protection intervenes.

  2. CPU, GPU, and SoC thermal zones heat and cool together through the shared package. This necessitates accounting for cross-domain heat propagation and retention when predicting future headroom.

  3. Hardware profiles trade off performance metrics: Hardware profiles with different CPU/GPU frequencies trade off time to first token (TTFT), time per output token (TPOT), and headroom consumption. This shows that minimizing power alone does not preserve serving performance, requiring a balance between TTFT, TPOT, and remaining thermal headroom.

System Architecture and Prediction Mechanism

ThermE separates the inference data path from the thermal control path. The core prediction mechanism relies on physics-informed neural networks (PINNs). ThermPINN identifies coupled thermal dynamics offline by minimizing a total loss function that includes terms for sensor temperature error, PDE residual, and boundary/initial-state conditions. This offline identification generates an orthonormal basis and reduced operators via Proper Orthogonal Decomposition (POD). The online Reduced Headroom Predictor (RHP) then advances the thermal states using these compact state updates during candidate evaluation.

Evaluation and Results

Implemented atop vLLM on an Nvidia Jetson AGX Orin, ThermE demonstrated significant performance improvements. Under the Stress workload at a 65°C ambient setpoint, ThermE reduced TTFT and TPOT by 40.55% and 12.17%, respectively, relative to vLLM, achieving a 5.70% SLO violation rate compared with 12.30% for the strongest baseline. The predictor obtained a 1.94◦C MAE with 24.36 ms overhead. Furthermore, the evaluation showed that ThermE lies on the non-dominated frontier in both serving quality and thermal exposure when comparing against baselines like vLLM and Static fixes, confirming its ability to balance serving progress with future headroom.

Ablation Studies

Ablation studies confirmed the necessity of each component. Removing the compiler, thermal coupling, or uncertainty margin resulted in increases in TTFT, TPOT, and SLO violations (11.02–14.32%). The Greedy scheduler increased TTFT and TPOT by 13.18% and 13.19%, demonstrating the benefit of anticipating how current actions consume future headroom. The full numerical solver resulted in the largest penalties, increasing TTFT by 8.70%, TPOT by 26.68%, and SLO violations by 48.20%. These results support combining predictive headroom estimation with multi-interval planning for sustained serving under thermal constraints.

Conclusion

ThermE provides a "predictive thermal resource management approach to the systems community, providing a foundation for sustained onboard intelligence in satellites, smart AI mobile devices, and other compact platforms like System-on-package (SoP) where limited cooling constrains continuous inference." It successfully coordinates hardware profiles and scheduled-token caps to preserve serving quality under changing workload and thermal conditions.

The gist: ThermE is a runtime system that predicts and jointly manages shared thermal headroom for sustained LLM inference on edge SoCs.

Improvements for AI systems

Based on the scientific paper ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs, here are specific, high-impact improvements that can be made to existing AI systems, and what those improved systems will be capable of:


The core improvement lies in moving from reactive hardware throttling to a proactive, holistic thermal resource management system. The proposed architecture allows AI systems deployed on edge SoCs (like autonomous vehicles, drones, or satellites) to operate at higher sustained performance levels without immediate service degradation due to thermal constraints.

Here are the specific improvements and their resulting capabilities:

The implementation of the compiler that maps request descriptors (input length, output budget) and runtime state directly to domain heat inputs for CPU, GPU, and RAM.

The development of a PDE-constrained Headroom Predictor utilizing ThermPINN for offline identification of coupled thermal dynamics and an online Reduced Headroom Predictor (RHP) for low-overhead uncertainty-calibrated forecasts.

The integration of an Uncertainty-Aware Action Scheduler that utilizes beam search to select optimal combinations of hardware profiles and scheduled token caps, explicitly balancing serving quality (TTFT/TPOT) against future thermal runway and risk of throttling.

These improved AI systems can perform the following specific functions:

Sustain High Throughput LLM Inference Under Extreme Thermal Stress:

The system can maintain significantly lower Time to First Token (TTFT) and Time Per Output Token (TPOT) compared to current reactive systems (like vLLM), even when operating near hardware throttling thresholds. The paper shows a potential reduction of up to 40.55% in TTFT and 12.17% in TPOT, meaning the system can keep generating tokens faster and more consistently during sustained periods (e.g., long context summarization or continuous dialogue).

Proactive Thermal Load Balancing Across Heterogeneous Components:

The system can intelligently adjust the operational frequencies of different processing units (CPU, GPU) and memory bandwidth dynamically. Instead of letting one component overheat and throttle everything, the improved system manages the shared thermal headroom to ensure that a performance-improving control decision does not prematurely consume shared resources needed for subsequent requests.

Adaptive Workload Scheduling Based on Predicted Thermal Futures:

The scheduler can make decisions based on a multi-interval (horizon) view of thermal state. It will select hardware profiles and token caps that are predicted to maintain a positive thermal runway over the next several control intervals, effectively saving headroom for future peak demands rather than consuming it instantly for the current request.

Robustness Against Environmental Fluctuations:

Because the predictor is trained on diverse trajectories (including varying ambient temperatures and cooling states), the system can accurately forecast how changes in external conditions (like entering direct sunlight or a hot environment) will affect internal thermal dynamics and available headroom, allowing it to preemptively adjust its operational strategy.

Optimized Power and Energy Consumption:

The scheduler explicitly incorporates domain energy costs into its scoring function, leading to better power distribution across the SoC. The improved system can achieve a superior trade-off between serving quality and energy efficiency compared to static or greedy approaches, allowing for longer operation times on battery-constrained edge devices.

Related papers