GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators

arXiv:2606.08343 · cs.LG · Submitted 2026-06-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators".

Jane: The paper was written by Jason Sulskis and Sathya Ravi from University of Illinois at Chicago and Georgia Tech Research Institute.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of GENERIC-FNO: Jane: So, we established that "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators" aims to fix a major flaw in current AI models which are essentially black boxes with no physical understanding.

Tom: Exactly, Jane. These existing neural operators just drift away from the physically admissible manifold over time because they don't know about conservation laws, so the authors show how they tackle this problem by embedding the full GENERIC structure into the function space itself.

Lu: It’s not just about adding a penalty term; it’s a fundamental architectural change that forces a deep thermodynamic consistency from Grmela and Öttinger's framework into the continuous field.

Meng: That sounds like they are building these systems using mathematical constraints rather than relying on the training data to learn physics, which is much more robust for large-scale deployment.

Lalam: The implication here is that we’ can finally have AI models that act like a virtual physics engine, respecting the laws of nature even when they are far away from where they were trained.

Tom: It’s a massive shift in thinking. Before Jane mentioned the specific mechanism, let's hear more details on how this works.

Improvements and Methodology: Jane: The authors have designed a method called "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators" that achieves degeneracy by construction, meaning it holds exactly to machine precision.

Tom: And what they are doing to achieve this is brilliant—instead of using a soft penalty, they parameterize the Poisson and friction operators as diagonal Fourier multipliers.

Lu: These multipliers are then sandwiched between rank-one projections which removes the conjugate gradient direction, which is how they enforce those crucial degeneracy conditions like L delta S / delta u = zero.

Meng: From an engineering viewpoint, this means that if they are running simulations using this AI model at all scales and resolutions, the underlying physics remains perfectly sound.

Lalam: The way they handle the gauge freedom is also very insightful; they separate what claims are invariant from what aren't, which is a very sophisticated way to talk about model attribution.

Tom: It’s really interesting how they achieve exactness without needing a penalty term, Jane. We've seen how this method performs in practice across different types of physical problems—reversible, dissipative, and mixed regimes.

Experimental Results and Performance: Jane: The results section for "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators" shows the performance across various operators and PDEs, comparing it to unconstrained baselines.

Tom: It's not just better on specific tasks; the structural guarantees hold zero-shot across a four times super-resolution range, which is an incredible feat for any AI model designed to generalize.

Lu: That zero-shot capability, combined with the exact structural preservation, means that we can trust the results even if the scale of our problem changes drastically.

Meng: The fact that it outperforms baselines on several dissipative and mixed problems while using comparable or fewer parameters is a very strong argument for efficiency and robustness in practical use.

Lalam: This suggests that adding this structural prior isn's just a theoretical curiosity, it actually yields tangible benefits in accuracy and stability across real-world scenarios.

Tom: The data speaks for itself, Jane. We’ve seen its performance, but let's discuss the implications of how it performs on the most challenging cases.

Conclusion and Final Thoughts: Jane: As we wrap up our discussion of "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators," it is clear that this paper has addressed a long-standing gap in the field of structural AI.

Tom: It' provides a neural operator with the full GENERIC structure, ensuring that energy is conserved and entropy is produced exactly by construction, which is a huge achievement.

Lu: The way they handle the gauge freedom and provide a falsifiable diagnostic method for physical dissipation is truly groundbreaking for the theoretical side of AI.

Meng: I think from an implementation standpoint, this will be a massive advantage in areas like fluid dynamics or weather modeling where maintaining physical plausibility is critical for making useful predictions.

Lalam: The final realization is that we' are building a new class of physical surrogates that truly respects the fundamental laws of nature, paving the way for much more reliable AI applications.

Tom: It’s been a fantastic discussion, Jane. We’ve covered the structure, the results, and what it will be to do with this paper in terms of future work.

Lu: I'm excited to see how this concept is applied to vector and multi-field systems next time around.

Meng: And I can't wait to see how it handles real-world industrial problems where stability is the only thing that matters.

Lalam: Let’s carry this idea forward into our next topic, knowing that the "GENERIC-FNO: Embedding Energy Conservation and Entropy Production into Fourier Neural Operators" provides a strong foundation for a more reliable future.

University of Illinois at Chicago · Georgia Tech Research Institute

cs.LG

Submitted: 2026-06-06

Updated: 2026-09-16

Importance score: 89/100

The gist: The paper introduces GENERIC-FNO, a novel framework designed to embed fundamental physical constraints—specifically energy conservation and entropy production—directly into Fourier Neural

Key concepts

GENERIC-FNO
This is a method that embeds the full GENERIC structure into the function space of Fourier Neural Operators. This architectural change forces deep thermodynamic consistency, allowing the AI model to respect physical laws, unlike existing black-box models.
Energy Conservation and Entropy Production
These are crucial physical constraints that the AI system must adhere to. The method enforces these rules by parameterizing operators as diagonal Fourier multipliers, ensuring the simulation remains physically sound across different regimes.
Zero-Shot Generalization
This is the model's ability to perform accurately without prior training for a specific task. The research demonstrates this capability across a four times super-resolution range, meaning its structural guarantees hold even if the scale of the problem changes drastically.

Terminology

Summary

The paper introduces GENERIC-FNO, a novel framework designed to embed fundamental physical constraints—specifically energy conservation and entropy production—directly into Fourier Neural Operators (FNOs). This methodology is critical for advancing scientific machine learning by ensuring that deep learning models respect the underlying laws of physics, thereby improving model robustness and generalization when simulating complex dynamical systems.

Methodology and Computational Overhead

The core of GENERIC-FNO involves enforcing physical identities by incorporating exact variational derivatives into the training process. The method requires computing delta E/delta u and delta S/delta u, which are achieved through two backward passes using reverse-mode autodiff, making this step intrinsic to the method rather than an implementation detail. While this overhead is substantial—the training step cost is reported to be approximately 10 times that of a plain FNO step—the authors note that the absolute step cost remains small (about 0.1 s). The computational complexity scales with the number of fields, suggesting future work on cheaper E, S gradients (shared backbones, a fused double-backward).

Validation of Physical Identities and Gauge Invariance

The model’s adherence to physical laws is rigorously validated across multiple metrics. Structural identities are confirmed to hold at machine precision; for instance, the reversible skewness delta E, L delta E and the two degeneracy conditions delta S, L delta E sit at or near the float64 machine epsilon. Furthermore, the framework distinguishes between gauge-dependent and gauge-invariant quantities. While the channel attribution for dissipative fraction rho M scatters wildly across seeds, the crucial diagnostic, r mech (the gauge-invariant dissipation rate), remains stable for a given flow. This stability confirms that thermodynamic claims rest on r mech.

Performance and Physical Ordering

The model demonstrates an ability to correctly capture physical ordering without explicit supervision on energy or entropy. When comparing the model's gauge-invariant rate r mech against the ground-truth rate, the results show that the reversible advection points cluster at the bottom-left (about 0) and the dissipative/mixed points lie up-and-to-the-right. This indicates that the model orders physical dissipation correctly without ever being supervised on energy or entropy. Additionally, while using explicit Euler for training is retained to ensure that the discrete update conserves energy to machine precision, the authors note that RK4 significantly reduces the per-step energy residual r E, raising it from about 10-7 (Euler) to about 10-3.

Evaluation and Diagnostics

Model performance is assessed using multiple diagnostic metrics on held-out trajectories. The primary measure of accuracy is the relative L2 error of a 10-step autoregressive rollout. Beyond this, the framework computes structural residuals (r E, degeneracy) and the gauge-invariant r mech. The learned functionals are shown to behave physically: for dissipative and mixed PDEs, the learned entropy S phi rises monotonically and tracks the decay of mechanical energy Q, while for reversible advection, both are flat (conserved).

Improvements for AI systems

The core innovation presented in this paper is the successful integration of exact variational calculus and fundamental physical conservation laws into deep learning architectures (FNOs), moving simulation beyond pure data fitting to physically constrained modeling.

Given the high stakes and complexity, my improvements focus on generalizing these principles from specific PDE domains (like fluid dynamics) to broader areas of scientific computing and inverse problems.


The Flaw Addressed: Current PINN/FNO architectures often treat physical constraints (E conservation, S increase) as simple loss terms (L physics = residual). This is mathematically weak and fails to guarantee structural integrity or gauge invariance.

The Improvement: Generalize the differentiable functional gradient machinery (delta E/delta u, delta S/delta u) into a modular, universal operator layer that can be applied to any field represented by a neural network output, regardless of whether it models fluid velocity, material stress, or biological concentration.

Technical Specification:

  1. Generalized Functional Input: The operator must accept not just the field u(x, t) and its time derivative d t u, but also a parameterized Lagrangian or Hamiltonian Density L[u, grad u, d t u].

  2. Automatic Variational Solver: Implement a dedicated module that performs the necessary symbolic differentiation (e.g., using frameworks like SymPy or specialized ML automatic differentiators) to compute the Euler-Lagrange equations and their corresponding variational derivatives delta L/delta u. This replaces the hardcoded E and S gradients with a general grad var.

  3. Constraint Enforcement: The training loss is then constructed not just from residuals, but from the minimization of structural deviations derived from these fundamental variational principles:

L Total = L Data + lambda E grad var, E squared + lambda S grad var, S squared

What the Improved System Can Do:

  • Robust Physical Simulation: Simulate any time-dependent physical system (e.g., electromagnetism, elasticity, reaction-diffusion) where a fundamental conserved quantity or governing Lagrangian is known, guaranteeing that the learned dynamics respect these laws to machine precision (like r E in Figure 6).

  • Model Selection and Validation: Automatically validate candidate physical models against the system's inherent symmetries and conservation laws before running expensive simulations, drastically reducing the search space for theoretical physics.

Sources

Related papers