A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics

arXiv:2608.02965 · cs.LG, cond-mat.mtrl-sci, cs.SY, eess.SY · Submitted 2026-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics".

Jane: The paper was written by Yachao Zhu, Qiujie Huang, Sinan Li, Yang Li, Gang Lei et al. from University of Technology Sydney and University of Sydney and National Railway Research and Design Institute of Signal and Communication.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a fresh preprint that just hit arXiv, and it's called "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics." Jane, I have to say, that title is a mouthful, but the ideas inside are genuinely exciting.

Jane: It really is a mouthful, Tom, but once you unpack it, it's about something we all rely on every single day. We're talking about the magnetic components inside power converters — the little transformers and inductors that manage electricity in everything from your phone charger to electric vehicle drivetrains. And this paper is trying to predict how those components behave when the power flowing through them changes rapidly.

Tom: Right, and that's the "transient magnetization" part. Normally, these components are designed assuming steady, smooth operation. But modern power electronics are switching faster and faster, and the magnetic materials inside are getting hit with all sorts of messy, non-sinusoidal waveforms. The old formulas just don't cut it anymore.

Jane: Exactly. The authors are from the University of Technology Sydney and the University of Sydney, and they've built a compact neural network that learns the relationship between the magnetic field and the flux density — the B-H curve — but in a way that respects the actual physics of hysteresis. That's the "physics-informed" part of the title.

Tom: And I love that they're not just throwing a giant black-box model at the problem. They're using a hybrid approach. There's a local branch that tracks the recent history and rate-dependent effects, and a global branch that looks at the whole waveform to understand the path-dependent memory of the material. It's like having both short-term memory and a big-picture view.

Jane: That's a great way to put it, Tom. The local branch is like remembering the last few steps you took, and the global branch is like seeing the whole map of where you've been. For magnetic materials, where you've been matters a lot — the material remembers its history, and that affects how it responds next.

Tom: And the results are pretty impressive. They tested it on fourteen different ferrite materials, and the model predicts the magnetic response with an average energy consistency error of just under two percent. That's with only about four thousand seven hundred trainable parameters per material — tiny compared to some of the baselines they compared against.

Jane: Small model, big accuracy. That's the kind of efficiency that could actually get adopted in real design workflows. But we're just scratching the surface here. I want to get into the actual methodology and why this physics-informed approach matters so much. Stick around, because next we're going to break down the problem they're solving and why it's so hard.

Tom: You heard Jane — we're just getting warmed up. Stay with us.

Summary: Jane: Welcome back. We're still on "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics," and Tom, I think we need to explain the core problem in a way that really lands for our listeners.

Tom: Please do, because I think the title alone might have scared some people off. What's the actual task here?

Jane: So imagine you have a magnetic core inside an inductor. You know the flux density — that's the B — because you can measure the voltage on a coil. But the magnetic field strength, the H, is what actually tells you about the current and the losses. The catch is, H doesn't just depend on the current B. It depends on the entire history of B. That's hysteresis.

Tom: And that's what makes it so tricky. If you've ever seen a B-H loop, it's not a single curve — it's a family of curves that depend on where you've been. So the paper frames this as a boundary-conditioned prediction problem. You give the model the measured history up to a certain point, then the future B trajectory, and it has to predict the H response.

Lu: If I can jump in here, Jane — that framing is what makes this paper stand out. It's not trying to predict a whole closed loop from scratch. It's predicting a segment of the trajectory given the past. That's much closer to what happens in a real converter, where the excitation is constantly changing and you never really reach a steady state.

Jane: Exactly, Lu. And the model architecture reflects that. The local branch is a gated recurrent unit that takes the recent B and H history and propagates the dynamic state forward. The global branch is a transformer that looks at the entire B waveform to capture the path-dependent hysteresis context. They fuse those two representations to predict H.

Tom: And they call it a "neural operator" because it's learning a mapping between functions — from the input B trajectory to the output H trajectory — rather than just a point-to-point mapping. That's a subtle but important distinction.

Meng: As someone who actually has to deploy these things, the parameter count is what catches my eye. four thousand seven hundred seventy-seven parameters per material. That's nothing. The baseline they compare against, MagLearn2, has over eight hundred sixty thousand parameters and still doesn't beat the energy consistency of this model. That's a huge win for practicality.

Jane: And the energy consistency is the key metric here. They're not just checking that the H waveform looks right point by point. They're checking that the accumulated work — the integral of H dB — matches the measured value. That's the quantity that relates to core loss, which is what designers actually care about.

Lu: Right, and they add a physics-informed regularization term during training that penalizes inconsistency in that accumulated energy. So the model is explicitly encouraged to produce trajectories that are not just accurate in shape but also energetically consistent. That's the "physics-informed" part really earning its keep.

Tom: So we've got a small model, a clever hybrid architecture, and a physics-based training objective. But I'm curious about what happens when you start pulling pieces out. That's where the real insight is, right?

Jane: Absolutely, Tom. The ablation study is next, and it tells us exactly which components matter and why. Let's get into that.

Improvements: Tom: We're back, still talking about "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics." Jane, you mentioned the ablation study — what did they find when they started removing parts of the model?

Jane: This is where the paper really earns its keep, Tom. They tested five different variations. First, they removed the global branch — the transformer that looks at the whole waveform. The sequence error jumped from about thirteen percent to nearly eighteen percent, and the energy error went up too. That tells you the full-waveform context is genuinely needed to capture hysteresis memory.

Tom: Makes sense. Without the big picture, the model is flying blind on path dependence. What about removing the local recurrent branch?

Jane: That one was even more dramatic. The peak-field error jumped to nearly eighteen percent, and the energy-sign mismatch rate — where the model predicts the wrong direction of energy flow — went up to over five percent. That's the branch that tracks the rate-dependent dynamic response, so without it, the model loses track of the actual field amplitude.

Lu: And that's the beauty of the hybrid design. The global branch captures the quasi-static hysteresis — the path dependence. The local branch captures the dynamic effects — the eddy currents and relaxation behavior. They're complementary, and the ablation proves it.

Meng: I was particularly interested in the ablation where they removed the incremental B input — the delta B. The sequence error only went up a little, but the energy error went up the most out of all the ablations. That's a really subtle finding.

Jane: Right, Meng. The instantaneous B tells you where you are on the flux density axis, but the delta B tells you which direction you're moving and how fast. Removing that hurts the accumulated energy consistency even though the pointwise waveform still looks okay. It's a great example of why you need multiple metrics to evaluate these models.

Tom: And the most surprising one to me was the last ablation, where they moved the temperature and start-position information from the initial hidden state to direct inputs. The sequence error ballooned to almost twenty-five percent. Same information, different placement, huge difference in performance.

Lu: That's a really interesting architectural insight. It suggests that conditioning the initial state is a much more effective way to inject context than feeding it as a stepwise input. The model needs that context to set up the boundary condition, not to influence every single step.

Jane: Exactly. And it shows that the authors thought carefully about where information flows in the network, not just what information is included. That's the kind of design discipline that makes a compact model work as well as much larger ones.

Tom: So what's the takeaway for someone designing power electronics? Is this ready for prime time?

Meng: I think the path is clear. The model is small enough to run on embedded hardware, it's trained per material, and it gives you both the waveform and the energy consistency. The next step would be integrating it into a circuit simulator to see how it performs in a full system context.

Jane: And that's exactly what the authors say in their future work — embedding this into electromagnetic and circuit-field simulation workflows. So we're not there yet, but the foundation is solid. Let's wrap up with our final thoughts.

Conclusion: Tom: And that brings us to the close of our discussion on "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics." Jane, give us the final word.

Jane: Tom, this paper is a great example of how thoughtful architecture design can beat brute-force scaling. They took a hard physics problem — predicting transient magnetic response under messy, real-world excitation — and solved it with a model that's tiny by modern standards. The hybrid of a recurrent local branch and a transformer-based global branch, guided by an energy-consistency loss, is a template that could apply well beyond magnetics.

Lu: I completely agree, Jane. The way they framed the problem as a boundary-conditioned trajectory mapping is elegant. It's not just about predicting a curve — it's about respecting the physics of memory and path dependence. That's a lesson that carries over to any system with hysteresis, from piezoelectric actuators to shape-memory alloys.

Meng: And from an engineering standpoint, the practical impact is clear. Better transient prediction means better core-loss estimation, which means smaller, more efficient magnetic components. That directly translates to smaller power supplies, better EV drivetrains, and more efficient data centers. The fact that it's compact enough to deploy is the cherry on top.

Tom: The authors also deserve credit for the thorough evaluation — fourteen different ferrite materials, multiple temperatures, multiple frequencies, and a rigorous ablation study. That's the kind of work that gives you confidence in the results.

Jane: Absolutely. And while the paper is a preprint, the methodology is sound and the results are compelling. We'll be watching to see where this goes, especially if they integrate it into circuit simulation tools as they suggest.

Tom: Well said. That's our time for "A Physics-Informed Hybrid Neural Operator for Transient Magnetization Prediction in Power Magnetics." Thanks to Lu and Meng for joining us, and to our listeners for tuning in. We'll be back soon with the next paper from the arXiv. Until then, keep asking questions and stay curious.

Jane: Take care, everyone. We'll see you next time.

Yachao Zhu, Qiujie Huang, Sinan Li, Yang Li, Gang Lei, Jianguo Zhu

University of Technology Sydney · University of Sydney · National Railway Research and Design Institute of Signal and Communication

cs.LG, cond-mat.mtrl-sci, cs.SY, eess.SY

Submitted: 2026-08-18

Updated: 2026-08-19

Comments: 13 pages, 7 figures. Preprint prepared for possible submission to IEEE Transactions on Power Electronics

Code: https://github.com/JunWang-Bristol/MagLearn2

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 57/100

Key concepts

Hysteresis
The phenomenon where a material's response depends on its history. In magnetic materials, the magnetic field strength depends on the entire path the material has taken, not just the current flux density. This 'memory' effect makes predicting magnetic behavior during rapid power changes complex.
Physics-Informed Training
This approach incorporates physical laws into the neural network's training process. By using an energy-consistency regularization term, the model is explicitly encouraged to produce trajectories that are not just accurate in shape, but also energetically consistent with the actual work performed by the magnetic material.
Hybrid Neural Operator
This architecture combines a local branch, using a gated recurrent unit to track recent history and rate-dependent effects, with a global branch, using a transformer to capture the entire waveform's context. This allows the model to understand both immediate dynamics and long-term path-dependent memory.

Terminology

Summary

Summary

This paper proposes the Physics-Informed Hybrid Neural Operator (PI-HNO), a compact, material-specific neural model designed for core-loss-oriented transient magnetization prediction in power magnetics. The work addresses the limitations of traditional steady-state core-loss formulas and single-valued material curves, which cannot fully capture transient magnetization responses under modern converter excitations involving fast transitions, minor-loop operation, dc bias, and temperature variation.

The prediction task is formulated as a boundary-conditioned trajectory mapping problem. Given a transient B–H sequence of L sampled time steps, a prediction boundary s divides the sequence into a measured-history interval (0 ≤ t < s) and a prediction interval (s ≤ t < L). The model receives the measured B(t)–H(t) history, the complete input B(t) trajectory over the prediction interval, temperature, and temporal-position information, and predicts the future H(t) response. The known-history ratio is defined as ρ = s/L, with test values of 0.1, 0.5, and 0.9.

The proposed architecture integrates two complementary learning branches motivated by physical mechanisms:

  • Local recurrent branch: A GRU encoder–decoder structure that learns boundary-conditioned evolution of rate-dependent magnetic response. The history GRU compresses measured B(t)–H(t) history into a hidden-state representation, which is conditioned by temperature and start-position projections. The future GRU propagates the dynamic representation over the prediction horizon using input B(t), its increment ΔB(t), and temporal-position embeddings.

  • Global attention branch: A Preisach-inspired Transformer-based branch that extracts waveform-level hysteresis context from the complete B(t) trajectory. Multi-head self-attention serves as a data-driven memory aggregation mechanism, analogous to the weighted superposition of elementary hysteresis operators in the Preisach model.

The model also incorporates an energy-aware training objective combining smooth-L1 loss for pointwise waveform matching with a soft B–H energy-consistency regularization term. The energy term penalizes inconsistency between predicted and measured accumulated B–H work over the prediction interval, computed via trapezoidal integration, plus a one-sided penalty for negative accumulated energy.

Experiments use the MagNetX transient database with material-specific models for 14 ferrite materials: 3C90, 3C92, 3C94, 3C95, 3E6, 3F4, 77, 78, N27, N30, N49, N87, FEC014, and T37. Excitation frequencies range from 50 to 800 kHz, and temperatures are 25, 50, and 70 °C. Sequences are fixed at L = 1000 time steps.

Key results:

  • PI-HNO achieves mean sequence error of 13.40% and 95th-percentile sequence error of 31.39%.

  • Mean B–H energy consistency error of 1.92% and 95th-percentile energy error of 7.60%.

  • Mean peak-field error of 6.12% and energy-sign mismatch rate of 1.91%.

  • The model uses only 4777 trainable parameters per material-specific model.

Comparisons with six baseline models (MagLearn2, LSTM, HARDCORE, MMINN, GRU-only, Transformer-only) show that while MagLearn2 achieves slightly lower sequence errors (12.88% mean, 28.47% 95th-percentile), PI-HNO achieves significantly better energy consistency (1.92% vs. 4.06% mean energy error) with approximately 1/180 of MagLearn2's parameter count. The GRU baseline with comparable parameters (4758) yields higher sequence errors (18.44% mean) and energy errors (3.68% mean).

Ablation studies evaluate five variants:

  • A1 (removing global branch): increases mean sequence error to 17.77% and mean energy error to 2.86%.

  • A2 (removing local recurrent branch): produces the largest peak-field error (17.80%) and energy-sign mismatch rate (5.57%).

  • A3 (removing ΔB input): gives the lowest sequence error among ablations but the largest energy errors (3.73% mean, 12.66% 95th-percentile).

  • A4 (removing future GRU): increases mean sequence error to 19.01% and 95th-percentile energy error to 12.18%.

  • A5 (bypassing temperature/start-position conditioning): gives the largest sequence degradation (24.83% mean, 75.68% 95th-percentile).

The paper concludes that PI-HNO provides a favorable trade-off among prediction accuracy, energy consistency, and model compactness, with future work focusing on embedding the model into transient electromagnetic and circuit-field simulation workflows for device-level loss prediction and thermal assessment.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to an AI system and what the improved system can do:

  • Integrate a dual-branch neural network combining a GRU-based local recurrent branch (for rate-dependent dynamic response) with a Transformer-based global attention branch (for Preisach-inspired hysteresis memory)

  • Implement boundary-conditioned prediction where the model receives measured B(t)–H(t) history before a prediction boundary, the complete future B(t) trajectory, temperature, and temporal-position features

  • Use additive latent-space fusion of local and global representations before the output projection, motivated by the quasi-static hysteresis + dynamic eddy-current decomposition

  • Add a soft B–H energy-consistency regularization term to the loss function that penalizes discrepancies between predicted and measured accumulated B–H work over the prediction interval

  • Include a one-sided penalty for negative accumulated energy to enforce physical directionality

  • Convert normalized B and H values back to physical units before computing the path integral using trapezoidal rule

  • Initialize the future GRU hidden state with a combination of history-GRU output, temperature projection, and start-position projection (rather than feeding these as stepwise inputs)

  • Include incremental flux-density input ΔB(t) alongside instantaneous B(t) to capture local excitation rate and direction

  • Use a complete B(t) trajectory input to the global branch to capture reversal points, flux-density extrema, dc offset, and minor-loop structure

  • Use only 4,777 trainable parameters per material-specific model (GRU width 8, Transformer width 8 with 2 heads, feed-forward width 64)

  • Apply dropout (0.1), gradient clipping (0.5), and exponential moving average (0.999) for stable training

  • Use one-cycle learning-rate scheduling with AdamW optimizer and peak learning rate of 1e-2

  • Predict future H(t) sequences from measured magnetic history, input B(t) trajectory, temperature, and temporal-position information

  • Reconstruct B–H trajectories over finite prediction intervals, including open trajectories (not just closed hysteresis loops)

  • Handle nonsinusoidal excitations with fast transitions, dc bias, minor-loop operation, and temperature variation

  • Achieve 1.92% mean and 7.60% 95th-percentile B–H energy consistency errors across 14 ferrite materials

  • Maintain low energy-sign mismatch rate of 1.91%, ensuring predicted trajectories preserve the direction of accumulated B–H work

  • Provide material-specific models that adapt to distinct hysteresis characteristics, dynamic behaviors, and temperature responses

  • Run with 1/180 of the parameters of the best baseline (MagLearn2) while achieving comparable or better energy consistency

  • Operate under varying known-history ratios (ρ = 0.1, 0.5, 0.9) to handle different levels of available magnetic-history information

  • Predict across frequencies (50–800 kHz) and temperatures (25–70 °C) with a single compact model per material

  • Distinguish prediction segments with similar instantaneous B(t) and ΔB(t) but different magnetization histories

  • Preserve rate-dependent behavior through sequential GRU state propagation along the excitation trajectory

  • Maintain accumulated-energy consistency even when pointwise sequence errors vary across materials, ensuring physically meaningful reconstructed trajectories

  • Demonstrate complementary contributions from each architectural component (global branch, local branch, ΔB input, future GRU, and conditioning strategy)

  • Achieve best overall balance of sequence accuracy, energy consistency, peak-field prediction, and energy-directionality compared to all baselines and ablations

Sources

Related papers