Is Your AI Fast Enough to Run a Fusion Reactor?
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Is Your AI Fast Enough to Run a Fusion Reactor?".
Dev: Machine learning models are increasingly used in feedback control loops for nuclear fusion, where inference speed and predictable timing are critical.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So this paper "Is Your AI Fast Enough to Run a Fusion Reactor?" really dives into how fast these machine learning models need to be for feedback control in fusion, which is something we always talk about. The authors are summarizing lessons from models used on the DIII-D tokamak and building a benchmark for different deployment backends, which seems super relevant given the real-time demands of fusion experiments.
Dev: I agree, Rosa; it seems like the central thesis here is that inference speed and predictable timing are absolutely critical in fusion control loops. The paper claims that for models larger than five million parameters, running them on a CPU backend takes tens to thousands of milliseconds, while using a GPU for inference is substantially faster, suggesting there's really an upper limit on what you can do with CPU-oriented development in this area.
Taro: That makes sense from an autonomy standpoint; if the response time is too slow, you lose the ability to react when things go sideways in a chaotic environment. The paper highlights that magnetic confinement control doesn't have a universal timescale because the physics itself is nonlinear and chaotic, meaning you have to run different tasks at different rates within one control cycle.
Rosa: Exactly, Taro; I was reading about how they show what must happen within one of these cycles, illustrating the whole pipeline from input diagnostics to actuator commands. It shows that input diagnostics like magnetics and cameras feed into signal processing, which then feeds the ML models, and finally control logic sends commands out to actuators.
Dev: And what’s really compelling is the different budgets they lay out for prediction warning times depending on the specific control task. They list things like needing a zero point eight ms inference time for beta N/ITB control within a fifty ms cycle, but then it drops to as low as zero point two ms for ELM prediction, which is needed for fast edge-event prediction. That variation really shows the precision required.
Taro: I wonder what happens when the world misbehaves and the response time is dictated by these tight budgets; for instance, if a tearing mode event could have been forecast about two hundred ms before its onset in a DIII-D tearing-mode experiment, can the system actually execute that fast prediction and response?
Rosa: That brings us to the benchmark they introduce, MILF-BENCH, which compares different deployment backends across ten neural networks and model components from various fusion control and diagnostic pipelines. This is where they test the practical performance under realistic conditions instead of just looking at theoretical speeds.
Paper summary: Dev: The MILF-BENCH protocol involves running inference with zero-copy streaming in C++ at a batch size of one for about ten seconds, with inputs requested every one thousand cycles. They compare Keras2c and OpenVINO running on the CPU against TensorRT FP16 and FP32 running on an NVIDIA Tesla V100S GPU for this comparison.
Taro: So, when we look at the model performance comparison, the results show that no single backend is fastest for every model. For instance, they show that TokaMind on a CPU using OpenVINO takes seventy-six point nine milliseconds to run inference, but using TensorRT FP32 on the V100S GPU gets it down to two point two zero milliseconds, which is a thirty-five times faster result.
Rosa: Wow, that difference really highlights the practical necessity of picking the right backend based on what you’re deploying. They showed that TensorRT provides the largest reductions for models larger than one million parameters, and in one comparison with INPA-Net, TensorRT FP16 was four point five times faster than the best CPU result.
Dev: It’s clear that for bigger models, GPU inference is the recommended path for control because of the significant latency reduction. However, they also pointed out that Keras2c has the lowest mean latency for smaller models up to ten thousand parameters due to its latency determinism.
Taro: That makes sense when you think about deployment; if your control loop needs very predictable timing, like for something running at thirty Hertz or one MHz, that deterministic performance might be a key factor in choosing Keras2c over something that fluctuates wildly.
Rosa: And the authors also made a note about quantization not having as strong an effect on latency as some people thought, suggesting keeping FP32 might be preferred over reducing to FP16 for certain applications. Plus, they stressed that jitter during inference is a critical measurement because lag spikes can happen right when critical physics phenomena are occurring.
Dev: That jitter point is huge for me; if the system has latency spikes, those spikes can coincide with dangerous plasma events, which means we need to measure that timing variation very carefully. So, we're not just looking at the average speed; we have to look at the worst-case scenarios too.
Taro: I think what really sticks with me is that for models bigger than five million parameters, they explicitly recommend considering GPU inference. That tells us that scaling up the model size pushes us out of comfortable CPU territory for reliable real-time control.
Paper summary: Rosa: So, to wrap up this discussion on "Is Your AI Fast Enough to Run a Fusion Reactor?", the paper is fundamentally about establishing how fast ML models need to be in fusion feedback loops and providing a benchmark for choosing the right backend. It moves us from just thinking about model accuracy to worrying intensely about deployment latency and timing consistency.
Dev: Indeed, it’s not just about achieving the lowest possible number; it’s about finding the right balance between speed and determinism for a control cycle that has such specific time budgets. The implications are that we have to tightly couple our model selection with the hardware we plan to deploy it on, especially when dealing with larger models exceeding five million parameters.
Taro: From a broader perspective, if this work helps us understand the practical latency limits for these complex feedback systems, it could inform how we design future autonomous systems that interact with highly dynamic physical environments. It shows us the constraints imposed by real-time physics on AI deployment.
Rosa: I think the title really captures the essence of what they’re doing; it’s not just about if the AI *can* run, but whether it's fast enough to maintain stability in a demanding physical system like fusion. It grounds the high-level discussion in very concrete timing requirements for those control loops.
Dev: And they clearly show that for anything big, we have to look at the GPU options because the speed difference is massive, like that thirty-five times difference they showed with TokaMind. It really solidifies the argument for GPU utilization when model size gets substantial.
Taro: So, as we look ahead to future work mentioned in the paper, I’m interested in how this benchmarking approach could be extended to even more complex control scenarios where the physics might evolve even faster than what’s currently modeled.
Rosa: That’s a good point, Taro; expanding the scope of these benchmarks to cover different timescales or physical phenomena would be a natural next step for this research direction. It shows that the current work sets a foundation for deeper investigation into real-time control constraints.
Dev: Exactly, and those future considerations are important because the limitations they mention—like Keras2c not being able to convert TokaMind because it lacks specific transformer modules—suggest there’s still room for model flexibility to be explored.
Paper summary: Taro: That limitation points toward the need for models that are inherently designed for these stringent control environments, rather than just being the largest possible general-purpose AI. The paper gives us a clear target for what kind of model architecture we need to pursue.
Rosa: It’s exciting to think about how this research could influence the development of future robotic systems that need to operate with such high temporal precision in unpredictable settings. It shows us the engineering reality behind deploying complex AI in demanding physical spaces.
Dev: I’m just focused on making sure that whatever we deploy, we measure those latency spikes and jitter rigorously because in a control loop, a timing hiccup can be fatal. That's the engineering reality we have to confront when designing these systems.
Taro: So, to summarize what we’ve heard from "Is Your AI Fast Enough to Run a Fusion Reactor?", the core message is that deployment backend selection must be tied directly to the model size and the required control cycle budget. The paper highlights that for large models, GPUs are necessary for speed, while maintaining deterministic timing is crucial for safety in these systems.
Rosa: That’s a solid overview of the main points from "Is Your AI Fast Enough to Run a Fusion Reactor?" and how it frames the critical timing challenges in fusion control with machine learning. We've seen how the authors use their benchmark to show that deployment choice matters more than just model capability alone.
Dev: And I think it really underscores that for any real-time control application, you can't ignore the physical constraints of the system; if the inference time doesn't fit within the cycle budget, the model is useless regardless of its accuracy. It’s a tight constraint we have to respect.
Taro: We should keep an eye on how this kind of benchmarking evolves as autonomy research pushes for more complex, real-world deployment scenarios where the environment is even more unpredictable. The paper opens up avenues for understanding these temporal constraints in broader fields.
Rosa: It’s clear that this paper provides a very practical guide for anyone working on deploying ML models in high-stakes, real-time control environments, whether it's fusion or something else. It’s about matching the right AI tool to the right physical reality.
Dev: And that practical guide is valuable because it forces us to confront the hardware realities upfront, rather than hoping a general-purpose framework will magically work under tight timing constraints. That's a necessary shift in how we approach this kind of engineering problem.
Conclusion: Rosa: So we've been looking at how these machine learning models need to run for feedback in fusion, and now we're getting to the big picture of this paper titled "Is Your AI Fast Enough to Run a Fusion Reactor?".
Dev: That title really hits the nail on the head because it’s all about making sure this AI can keep up with the physical reality of a reactor.
Taro: I think what this paper is really getting at is that we're moving past just asking if an AI can be accurate enough, and starting to ask if it has the timing to actually make decisions when things get chaotic.
Rosa: Exactly, and the authors are showing us through their testing that speed and predictability are just as important as the accuracy itself for real-time control.
Dev: I agree; I mean in control engineering, a slow loop rate means you’re operating on old information, which is a serious failure mode when dealing with plasma instabilities.
Taro: And that's where the autonomy question comes in; if the system lags by even a fraction of a second during an event, what happens when we need to react instantly?
Rosa: Well, this paper lays out exactly how fast those timing constraints are for different control tasks within fusion experiments.
Dev: It really drills down into the specific millisecond budgets required for things like ELM prediction or tearing mode avoidance, which gives us a concrete performance target.
Taro: That focus on specific timescales is crucial because the physics itself operates on different speeds, so we need models that can handle those varying rates.
Rosa: So, it seems this research is essentially giving us the engineering roadmap for deploying AI in these highly demanding physical environments.
Dev: It’s a practical guide showing us how to select the right hardware and model architecture to meet those tight timing requirements on the ground.
Taro: And that selection process isn't just about picking a fast algorithm; it’s about understanding where you're hitting your hard limits based on the physics of the system.
Rosa: It really shows how deeply integrated AI has to be with the physical constraints of something like a fusion reactor for it to actually function reliably.
Dev: And that leads us right into how this kind of performance testing translates into real-world deployment scenarios outside of controlled lab settings.
Nathaniel Chen, Andrew Rothstein, Ricardo Shousha, Hiro Farre-Kaga, Peter Steiner, Azarakhsh Jalalvand, Egemen Kolemen
Princeton University
physics.plasm-ph, cs.PF, cs.SY, eess.SY
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/NVIDIA/TensorRT
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Machine learning models are increasingly used in feedback control loops for nuclear fusion, where inference speed and predictable timing are critical.
Key concepts
- Fusion Timescales
- The physical processes in a fusion reactor occur at vastly different speeds. Some events, like vertical displacement, happen in milliseconds, while others, such as temperature changes or density profiles, evolve over hundreds of milliseconds. This means control systems must handle tasks requiring very fast responses alongside those needing slower updates.
- MILF-BENCH
- This is a benchmark protocol created to compare how quickly different machine learning models run on various hardware. It tests ten neural networks and model parts using zero-copy streaming in C++ at high input rates, helping researchers decide if a model is fast enough for the reactor's control needs.
- Inference Latency
- This refers to the time it takes for a machine learning model to process input data and produce an output command. In fusion control, this speed is critical because lag spikes can disrupt real-time physical processes, making low latency essential for safe operation.
Terminology
Summary
Machine learning models are increasingly used in feedback control loops for nuclear fusion, where inference speed and predictable timing are critical.
The central processing unit (CPU) backends take tens to thousands of milliseconds for models greater than five million parameters, while graphics processing unit (GPU) inference is substantially faster, suggesting an upper limit on CPU-oriented development for control.
Fusion Timescales Needed for Inference
Magnetic confinement control does not have a universal timescale because much of the physics that occurs during fusion is nonlinear and chaotic. On DIII-D, physically disastrous phenomena such as vertical displacement events can evolve on the order of milliseconds, while slower profiles such as temperature and density evolve over more than 100 ms. This necessitates running different tasks at different rates within a control cycle.
Figure 1 illustrates what must happen within one of these cycles: Input diagnostics include magnetics, interferometry, electron cyclotron emission, beam emission spectroscopy, charge exchange recombination, Thomson scattering, and cameras. Signal processing converts these measurements into quantities such as a magnetic equilibrium reconstructed in real time. ML models consume the resulting signals, and control logic converts their outputs into commands for actuators not limited to neutral beams or gas valves.
Budgets for Prediction Warning Times
The required inference speed varies significantly depending on the control task. Table 1 details these timing constraints:
-
βN /ITB control requires a 0.8 ms inference time within a 50 ms cycle, driven by slower profile evolution and NBI/ECH power adjustments.
-
Tearing-mode avoidance needs less than 10 ms, requiring an update within a 200 ms warning horizon for gyrotron steering.
-
ELM prediction requires 0.2 ms inference time to feed avoidance logic for fast edge-event prediction.
-
Alfvén-eigenmode control demands 0.65 ms inference time for fast fluctuation response via NBI power adjustments, and Profile control requires 15 ms for profiles that evolve over more than 100 ms using multiple actuators.
Benchmarking Deployment Latency (MILF-BENCH)
The paper introduces the Machine-learning Inference Latency for Fusion BENCHmark (MILF-BENCH) to compare deployment backends across ten neural networks and model components from fusion control and diagnostic pipelines. The benchmark compares Keras2c, OpenVINO, TensorRT FP16, and TensorRT FP32 across various workloads.
The protocol involves running inference with zero-copy streaming in C++ at batch size 1 for approximately 10 seconds with inputs requested at 1 kHz. Slower configurations complete fewer calls. The comparison is structured around:
: Keras2c and OpenVINO run on an Intel Xeon Platinum 8268 CPU at 2.90 GHz. TensorRT runs in FP16 and FP32 on an NVIDIA Tesla V100S PCIe 32 GB GPU.
Model Performance Comparison
The results show that No backend is fastest for every model.
Keras2c has the lowest mean latency for one model, OpenVINO for five, TensorRT FP16 for two, and TensorRT FP32 for two. The choice of backend depends on the specific model being deployed.
TensorRT provides the largest reductions for models beyond 1M parameters. For example, in comparing INPA-Net:
: Mean latency is 2.28 ms with Keras2c, 0.311 ms with OpenVINO, 0.0693 ms with TensorRT FP16, and 0.0700 ms with TensorRT FP32.
This means TensorRT FP16 is 4.5× faster than the best CPU result.
For TokaMind, OpenVINO on the CPU takes 76.9 ms, whereas TensorRT FP32 on the V100S takes 2.20 ms, a 35× difference.
Conclusion and Recommendation
The paper concludes that for models greater than five million parameters, it is recommended to consider GPU for inference.
While Keras2C is advantageous when running inference up to 10k parameters due to its latency determinism, larger models justify GPU inference. Furthermore, the authors note that quantization does not seem to have as strong an effect on latency; therefore, keeping FP32 may be preferred over floating point reduction to FP16.
Jitter during inference is also a critical measurement for real-time control because lag spikes can coincide with critical physics phenomena.
Limitations
The benchmark has limitations regarding model compatibility. Keras2c cannot convert TokaMind because it does not implement the transformer modules that the model requires.
Additionally, TensorRT FP16 was not measured for TokaMind because portions of the model need to be additionally quantized to function correctly.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems, based on this research, and what those improved systems can achieve:
-
Improve deployment strategy by selecting a backend tailored to the model's parameter count and required latency budget.
-
Implement model-specific inference optimizations using specialized runtimes (TensorRT for GPU acceleration, OpenVINO for optimized CPU execution) to drastically reduce latency compared to standard Keras2c or unoptimized CPU backends.
-
Develop robust control systems capable of handling asynchronous inference by utilizing the last completed output during delays, ensuring continuous command issuance even when inference lags.
-
Design safety checks that are decoupled from the learned model, allowing for easy testing and limiting the impact of a specific model failure on overall system stability.
-
Employ advanced techniques like FP16 quantization (or careful selection of FP32) to balance inference speed gains with acceptable latency, as it provides significant speedup for large models like INPA-Net and TokaMind.
-
Implement deterministic execution methods (e.g., CPU affinity, pre-allocated buffers, synchronous execution where appropriate) to minimize timing variation (jitter), which is critical when inference spikes coincide with critical physics events.
-
For models exceeding five million parameters, prioritize GPU inference using TensorRT configurations due to the substantial latency reduction observed compared to CPU methods.
These improvements result in AI systems that can:
-
Maintain real-time control performance on demanding physical systems (like fusion reactors) by ensuring inference times fit within strict, mission-critical control cycle budgets (e.g., sub-millisecond responses for tearing mode avoidance).
-
Predict plasma events with greater lead time and accuracy by deploying models optimized for the specific data pipeline they operate in.
-
Operate reliably under fluctuating input rates (1 Hz to 2 MHz diagnostics) by managing latency spikes without causing control system instability or blocking necessary actuator commands.
-
Handle extremely large, complex models (like multimodal transformers) that were previously infeasible for real-time deployment by leveraging high-performance GPU hardware and precision reduction techniques.
Abstract
Machine learning models are increasingly used in feedback control loops for nuclear fusion, where inference speed and predictable timing are critical. We summarize lessons from models deployed for control on the DIII-D tokamak and develop a benchmark to compare inference backends across ten neural networks and model components from fusion control and diagnostic pipelines. For models greater than five million parameters, the CPU backends take tens to thousands of milliseconds, while GPU inference is substantially faster, suggesting an upper limit on CPU-oriented development for control. These results show why the deployment backend must be selected together with the model and its control-cycle budget.
Sources
- TokaMind: A Multi-Modal Transformer Foundation Model for Tokamak Plasma Dynamics
- Towards Large-Scale Heterogeneous Data Organization for Scientific Foundation Models: A Nuclear Fusion Case Study
- TokEye: Fast Signal Extraction for Fluctuating Time Series via Offline Self-Supervised Learning From Fusion Diagnostics to Bioacoustics
- FPGA-Accelerated Real-Time Diagnostics at DIII-D Using the SLAC Neural Network Library for ML Inference
- Robust Control of ECH Deposition Profiles on DIII-D
- TorbeamNN: Machine learning based steering of ECH mirrors on KSTAR
Related papers
- GyroSwin: 5D Surrogates for Gyrokinetic Plasma Turbulence Simulations
- Electromagnetic ghosts in pair plasmas
- Collisionless whistler heat-flux instability in ultra-high- beta plasmas
- Rugged magneto-hydrodynamic invariants in weakly collisional plasma turbulence: Two-dimensional hybrid simulation results
- Hybrid Fourier Neural Operator-Plasma Fluid Model for Fast and Accurate Multiscale Simulations of High Power Microwave Breakdown
- Experimental evidence for coronal mass ejection suppression in strong stellar magnetic fields