Virtual Temperature Sensors in Power Transformers Using Neural Ordinary Differential Equations
Berk Hadzhamolla, Alexander Johannes Stasik, Signe Riemer-Sørensen
University of Oslo · SINTEF AS · Norwegian University of Life Sciences
cs.LG
Submitted: 2026-08-13
Updated: 2026-08-14
Journal ref: PHM Society European Conference, 9(1), 1-9, 2026
DOI: 10.36001/phme.2026.v9i1.4991
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper presents a physics-aware Neural Ordinary Differential Equations (Neural ODEs) framework for virtual temperature sensing in power transformers, applied to real-world time-series data from
Terminology
Summary
This paper presents a physics-aware Neural Ordinary Differential Equations (Neural ODEs) framework for virtual temperature sensing in power transformers, applied to real-world time-series data from fifteen transformers in the Norwegian transmission grid. The work addresses limitations of existing approaches: high-fidelity numerical methods (FEM, CFD) are computationally prohibitive and require unknown geometries; lumped-parameter models depend on transformer-specific constants; and purely data-driven ML methods (ANNs, LSTMs, CNNs) require large datasets and risk physically inconsistent results.
The methodology encodes the continuous-time heat-transfer structure of the transformer system directly into the Neural ODE architecture. The thermal dynamics are modeled with coupled energy balance equations for winding and oil temperatures, where the rate of change of each thermal state depends only on current states and external inputs (first-order Markovian structure). The state vector y(t) contains oil temperature, high- and low-voltage winding temperatures, and high- and low-voltage hotspot temperatures; the control vector x(t) contains high- and low-voltage active power loads and ambient temperature. The physics-awareness is embedded in two ways: (1) inputs to the neural network at each integration step are exactly (y(t), x(t)), preserving the physical decomposition between internal states and external excitations; (2) the network always outputs dy/dt rather than predicting temperatures directly, enforcing continuous-time heat-transfer structure. Exogenous inputs are provided at every integration step via piecewise-linear interpolation of discrete 15-minute observations.
The vector field is a three-layer fully connected network with hidden dimension 256 and tanh activations, trained with a composite loss combining trajectory mean squared error and a smoothness penalty (λ = 10−4) that encodes the physical prior that temperatures cannot change rapidly. Training uses the adaptive dopri5 ODE solver (fixed-step rk4 was unstable), Adam optimizer with learning rate 10−4, weight decay 10−5, gradient clipping norm 1.0, and a curriculum learning strategy where the prediction horizon grows linearly from 5 steps (1.25 h) to 192 steps (2 days) over training epochs. A three-layer stacked LSTM baseline (hidden dimension 256, dropout 0.1, scheduled sampling) is used for comparison with identical optimizer settings and information constraints.
Results show three performance regimes. In the unstable regime (training ≤ 1 month), with one week of data median R2 = −3.29 and MAE = 4.55 °C, indicating catastrophic trajectory divergence; one month improves to R2 = −0.75 but remains below useful prediction. From three months of training onward, the model enters a stable and useful regime: median R2 reaches +0.387 at three months and improves monotonically to 0.521 at one year, while median MAE drops from 2.07 °C to 1.73 °C. Beyond one year, performance plateaus at approximately R2 = 0.59 for forecast horizons of 30 days and longer, with median MAE stabilizing around 1.96–1.99 °C. The three-month threshold has a natural physical interpretation: three months of data at 15-minute resolution amounts to 8,640 samples spanning a full seasonal quarter, and transformer thermal behavior is strongly coupled to ambient temperature which varies by up to 30 °C between winter and summer in Norway.
In head-to-head comparison across all 823 evaluated combinations of unit, training horizon, and forecast horizon, the Neural ODE achieves higher R2 in 508 cases (61.7%) and lower MAE in 533 cases (64.8%). In the operationally relevant stable regime (training ≥ 3 months, forecast ≥ 30 days), these figures rise to 65.2% on R2 and 70.1% on MAE, with median R2 = 0.591 versus 0.493 for the LSTM and median MAE = 2.04 °C versus 2.45 °C – a difference of 0.41 °C. The Neural ODE is the better model on MAE at every forecast horizon without exception. At three months of training, the Neural ODE achieves median R2 = +0.39 and MAE = 2.07 °C while the LSTM median R2 is effectively zero, with best performance rates reaching 84% on R2 and 86% on MAE – the highest of any training horizon. At one year, NODE median R2 = 0.52 versus LSTM = 0.37; at two years, the LSTM narrowly wins on median R2 (0.52 vs. 0.38) though the Neural ODE retains lower MAE throughout.
Per-unit analysis in the stable regime shows the Neural ODE attains higher median R2 on 14 of 15 units and lower median MAE on all 15 units. The highest Neural ODE median R2 values are obtained for U01 (0.909), U09 (0.825), and U15 (0.773). Units 05 and 06 remain the most challenging for both models, with the Neural ODE still attaining positive median R2 (0.214 and 0.159) while the LSTM yields negative values (−0.092 and −0.148). The only unit where the LSTM achieves higher median R2 is U14 (0.699 versus 0.475), though the Neural ODE retains slightly lower median MAE (1.90 °C versus 1.94 °C). Qualitative comparisons (Figures 4 and 5) show the Neural ODE tracks measured trajectories more closely, while the LSTM diverges progressively during low-temperature periods.
The paper concludes that the Neural ODE consistently outperforms the LSTM baseline in the operationally relevant regime, with the most significant practical advantage being data efficiency: the Neural ODE produces useful forecasts from as little as three months of training data, whereas the LSTM median R2 remains near zero in this regime. The main limitation is a minimum data requirement below roughly three months, where learned dynamics become unstable. Future work will investigate warmup strategies, regularization methods, multi-unit transfer learning, probabilistic extensions, and online adaptation.
Improvements for AI systems
Improvements to AI Systems:
-
Physics-encoded continuous-time dynamics: Replace discrete-step recurrent architectures (e.g., LSTMs) with Neural ODEs that output time-derivatives of physical states, enforcing first-order Markovian heat-transfer structure. This reduces data requirements by 75% (useful from 3 months vs. >1 year for LSTM) and improves forecast MAE by 0.41 °C in stable regimes.
-
Smoothness-constrained training: Add a physical prior penalty (λ=10−4) on the derivative magnitude to prevent unphysical rapid temperature fluctuations, improving trajectory stability and reducing divergence in low-data regimes.
-
Curriculum learning on prediction horizon: Linearly grow forecast horizon from 5 to 192 steps during training, enabling the model to learn short-term dynamics first before long-term extrapolation—this stabilizes training and prevents catastrophic divergence.
-
Exogenous input interpolation: Provide piecewise-linear interpolated external drivers (load, ambient temperature) at every ODE integration step, allowing continuous-time modeling with discrete 15-minute sensor data without information loss.
-
Adaptive ODE solver selection: Use dopri5 (adaptive step-size) instead of fixed-step rk4, which was unstable, ensuring numerical robustness during training and inference.
What the improved AI system can do:
-
Predict transformer hotspot and oil temperatures up to 30+ days ahead with median R2=0.59 and MAE≈2.0 °C, using only 3 months of 15-minute historical data—a regime where LSTM-based systems fail (R2≈0).
-
Operate reliably across heterogeneous units (15 transformers) without per-unit recalibration, achieving positive R2 on all units and outperforming LSTM on 14/15 units.
-
Maintain physical consistency (no rapid temperature jumps) and generalize across seasonal variations (30 °C ambient swings) due to embedded heat-transfer structure.
-
Provide accurate forecasts even with sparse training data (e.g., one quarter), enabling deployment in new installations or regions with limited historical records.
-
Achieve lower error (MAE) at every forecast horizon compared to pure data-driven models, making it suitable for predictive maintenance and grid-load management in energy systems.
Abstract
Accurate modeling and forecasting of power transformer thermal behavior are critical for reliability, asset lifetime, and optimized power system operation. Numerical approaches such as finite element methods (FEM) and computational fluid dynamics (CFD) offer high fidelity but are computationally expensive, require complex mesh generation, and are often impractical for real-time or large-scale applications, particularly when transformer geometries are unknown. Lumped-parameter thermal models are more practical but depend on transformer-specific thermal constants and may fail to capture dynamic responses under varying operating and environmental conditions. Purely data-driven machine learning methods, including artificial neural networks, convolutional neural networks, and long short-term memory (LSTM) networks, have shown success in forecasting transformer temperatures but typically require large volumes of high-quality training data and may produce physically inconsistent or uninterpretable results. This paper develops a physics-aware Neural Ordinary Differential Equation (Neural ODE) framework for forecasting transformer thermal behavior from real-world time-series data. Neural ODEs model system dynamics in continuous time, providing smooth trajectory prediction and a natural representation of continuously evolving thermal dynamics. A key contribution is the integration of simplified heat-transfer equations directly into the Neural ODE formulation. The model is evaluated across datasets from fifteen transformers in different regions of Norway with varying designs and cooling mechanisms. The results demonstrate that the developed Neural ODE framework provides a standardized, physics-aware, and robust forecasting approach for heterogeneous transformer units.
Sources
- Neural Ordinary Differential Equations
- On Neural Differential Equations
- Neural Controlled Differential Equations for Irregular Time Series
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks