Learning to Emulate Chaos: Adversarial Optimal Transport Regularization

arXiv:2604.21097 · stat.ML, cs.LG · Submitted 2026-04-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Learning to Emulate Chaos".

Jane: The gist The authors propose a family of adversarial optimal transport objectives to jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory,

Tom: First, who's behind it and why it matters.

Title and authors: Jane: Let’s talk about the title and who wrote this. "Learning to Emulate Chaos: Adversarial Optimal Transport Regularization." It tells you immediately that they are using a specific mathematical tool, optimal transport, to train an AI emulator for chaotic systems.

Tom: Yeah, the title makes it clear that this isn't just another loss function tacked onto a standard model; it’s about fundamentally changing how the model learns the underlying physics of chaos by optimizing those transport costs #pg1.

Lu: Gabriel Melo and Leonardo Santiago are driving this work, and they’ve introduced these new statistics-based losses for emulating chaos #pg2. It's about making sure the emulator learns what matters statistically, not just point-to-point accuracy.

Meng: So, if we distill that down for someone who doesn't know math too well, it means they are training an AI to learn the essential features of a chaotic system from just one recording instead of needing massive datasets.

Jane: Right. It’s about efficiency and robustness, allowing the AI to learn the essential structure of complex attractors without getting overwhelmed by noise in a single observation #pg1.

Tom: And this is crucial because traditional methods struggle with long-term forecasts because of that sensitivity to initial conditions #pg1. This paper suggests a way around that fundamental difficulty.

The paper's summary: Jane: So, what’s the core idea behind this whole approach? It boils down to training an emulator and a set of summary statistics simultaneously, where the statistics are learned adversarially to match the data distribution #pg2.

Tom: That adversarial setup is key because it ensures the AI doesn't just learn any random features; it learns a specific set of statistics that are optimally informative for describing that chaotic attractor #pg2.

Lu: They combine a standard mean squared error loss for one-step prediction with this optimal transport cost to make sure the model’s summary statistics match the true data distribution #pg4.

Meng: I see how they are trying to solve that trade-off between being accurate step-by-step and being statistically consistent over time #pg2.

Jane: They do it by training the summary map to maximize that transport cost, which forces it to find the most discriminative statistics possible #pg2.

Tom: And they explore two ways this can be done: a WGAN-style dual formulation for p equals one, and a Sinkhorn divergence approach for any p greater than or equal to one #pg4.

Lu: The sinkhorn divergence is particularly useful because it’s fully differentiable, which means the AI can train it without worrying about those tricky Lipschitz constraints #pg8.

The paper's improvements: Tom: Let’s look at what they actually improve over existing methods. They point out that this loss function is much more robust to noise compared to standard mean squared error training #pg2.

Jane: That robustness is tied to the analysis showing that the optimal transport regularizer is controlled by the one-step prediction error, which implies local stability in the model's behavior #pg5.

Lu: They also show that for k-step rollouts with noisy initial conditions, the MSE loss scales in a way that suggests long-term consistency even when things get noisy #pg6.

Meng: This is interesting because it hints at a horizon-dependent training strategy; they can down-weight the standard prediction error after a certain mixing time if the MSE term becomes dominated by noise #pg4.

Tom: And beyond just matching attractor statistics, this method also allows them to estimate other dynamical properties like the leading Lyapunov exponent, or LLE #pg10.

Jane: So for someone just listening to the show, it means this AI can give us a better sense of how chaotic the system is—how fast things are diverging—which is a very useful physical property #pg10.

Conclusion: Tom: So, wrapping up the paper "Learning to Emulate Chaos: Adversarial Optimal Transport Regularization," this work gives us a framework where we can learn high-quality statistics and a consistent emulator from just one noisy trajectory #pg2.

Jane: The big implication is that this distributional regularization prevents the attractor from collapsing under noise, which is something standard mean squared error training often fails to do over long rollouts #pg15.

Lu: It’s a powerful way to ensure that whatever the emulator learns, it's statistically grounded in the true chaotic dynamics, even when you only have noisy data #pg2.

Meng: From my view as an engineer, the practical win is getting a much more reliable long-term prediction signal when you’re dealing with real-world data that isn't perfectly clean #pg15.

Tom: Exactly. It confirms that enforcing the statistics of complex, high-dimensional chaotic attractors via this adversarial optimal transport objective provides a useful signal toward the invariant measure for long-term evaluation #pg15.

Jane: It’s a solid piece of research that shows how we can learn what matters about chaos by letting the system figure it out through this specific mathematical structure #pg15.

Institut Poly-technique de Paris · North Carolina State University · Tufts University

stat.ML, cs.LG

Submitted: 2026-04-22

Updated: 2026-10-08

Code: https://github.com/gabrielmelo00/LearningToEmulateChaos_ICML26

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 82/100

The gist: The gist The authors propose a family of adversarial optimal transport objectives to jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory,

Key concepts

Adversarial Optimal Transport Objectives
This is a loss function strategy where two components—an emulator and a summary map—are trained against each other. The goal is to make the emulator's outputs statistically match the true data by minimizing an optimal transport cost, while simultaneously training the summary map to maximize that same cost, ensuring the learned statistics are highly informative.
Optimal Transport Cost (Wasserstein Term)
This mathematical term measures the minimum 'cost' of transforming one probability distribution into another. In this context, it forces the emulator's learned statistics to have a distribution that is optimally close to the true data's distribution, providing a strong constraint on how well the model captures the system's underlying statistical structure.
Sinkhorn Divergence
This is a specific mathematical relaxation of optimal transport used in the proposed loss function. It allows for stable, fully differentiable gradients without needing complex constraints like Lipschitz conditions on a critic. It provides an entropy-regularized way to enforce the desired distribution matching between the model and real data.
Long-Term Statistical Fidelity
This refers to how well an emulator can accurately predict the statistical properties of a chaotic system over many time steps, especially when starting from noisy initial conditions. The paper shows that this approach significantly improves fidelity compared to standard methods, helping the model capture the true invariant measure of the attractor.

Terminology

Summary

The gist The authors propose a family of adversarial optimal transport objectives to jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory, which significantly improves long-term statistical fidelity for emulating chaotic systems.

How it works

  1. The core idea is to train an emulator that learns the dynamics of the chaotic system while simultaneously learning an adversarial set of informative summary statistics, enforcing their distribution via an optimal transport cost on the emulator (Page 2). This involves training both the emulator and a summary map, where the latter is trained to maximize the same optimal transport cost (Page 2).

  2. The objective function combines a standard Mean Squared Error (MSE) loss for one-step fidelity with an optimal transport cost that matches the summary statistic distribution of the model to the data, while training the summary map to maximize this cost (Page 4). This is formulated as: min g∈G max f∈F L(g, f):= LMSEut+1, g(ut) + λ WdS p f%mu, f%muˆ (Page 4).

  3. The Wasserstein term can be implemented in two main ways:

- WGAN-style dual formulation:

When p = 1, the Wasserstein term is reformulated using the Kantorovich–Rubinstein duality, leading to a min-max objective where a critic φ acts as a statistical witness distinguishing the pushforward invariant measures (Page 4). This formulation is differentiable almost everywhere and can be optimized using stochastic gradient methods with standard Lipschitz-enforcing heuristics (Page 4).

- Sinkhorn divergence:

The authors adopt a primal, entropically regularized optimal transport objective based on the Sinkhorn divergence, which provides stable, fully differentiable gradients without requiring explicit Lipschitz constraints on a critic (Page 4). This is defined as: min g∈G max f∈F LS(g, f):= LMSEut+1, g(ut) + λ Sε cf%mu, %muˆ (Page 8).

Theoretical Analysis and Robustness

The theoretical analysis establishes that the proposed loss is well-behaved and provides an efficient way to capture the structure of high-dimensional chaotic attractors (Page 2). Specifically, Proposition 5.1 shows that the optimal transport regularizer is controlled by the one-step prediction error, implying local stability (Page 5). Furthermore, Corollary 5.9 demonstrates that for k-step rollouts with noisy initial conditions, the MSE satisfies L(k),noisy MSE ≳ dσ21 + σ22 · d · e2λmaxk as k → ∞ (Page 6).

The Wasserstein component exhibits robustness to noise; Theorem 5.14 shows that the Wasserstein distance stabilizes to a finite value determined only by measurement noise and model error in preserving the invariant measure, meaning lim k→∞ Wp f%muobs, f%muˆnoisy k ≤ Lf σ1κp,d + Wp%mu, ν (Page 7). This suggests a horizon-dependent training strategy where the Wasserstein term provides a useful signal toward the invariant measure when MSE becomes noisedominated at long horizons (Page 7).

Empirical Validation and Results

Numerical experiments across three high-dimensional chaotic systems—Lorenz-96, Kuramoto-Sivashinsky, and Kolmogorov flow—show that emulators trained using the proposed objectives have significantly improved long-term statistical fidelity compared to standard MSE training or prior statistics-based approaches (Page 3).

- Performance Comparison:

Table 2 reports that for L96 (multi-traj), the WGAN method achieves a lower L1 histogram error of 0.175 under noisy conditions compared to the No OT baseline of 0.348 (Page 9). For KS (single-traj), the WGAN method achieves an L1 histogram error of 0.435, outperforming the Fixed OT baseline which is at 0.290 (Page 9).

- Dynamical Property Improvement:

Beyond attractor statistics, the approach improves dynamical properties such as the leading Lyapunov exponent (LLE), which represents the rate at which chaos becomes unpredictable (Page 10). For Lorenz-96, our WGAN-style approach produces an emulator with an LLE of 2.336, nearly matching the ground truth value of 2.334, while the No OT baseline and Fixed OT fall significantly short (Page 10).

- Geometric Fidelity:

Figure 8 shows space-time plots for Lorenz-96, where the Learnable WGAN most faithfully replicates the diagonal wave patterns and spatial structure of the ground truth across the full evaluation rollout (Page 25).

Conclusion

The authors have proposed a family of adversarial optimal transport objectives that jointly learn high-quality summary statistics and a physically consistent emulator from a single chaotic trajectory, demonstrating their effectiveness in enforcing the statistics of complex, high-dimensional chaotic attractors (Page 8). This approach automatically discovers informative statistics without requiring domain expertise and can learn from a single noisy trajectory (Page 8). The results confirm that distributional regularization prevents attractor collapse under noise and provides a useful signal toward the invariant measure for long-term evaluation in chaotic systems (Page 15). Future work will explore extending this framework to stochastic systems by capturing the interplay between stochasticity and chaotic dynamics (Page 9).

--- Page 1 ---

Learning to Emulate Chaos: Adversarial Optimal Transport Regularization Gabriel Melo Leonardo Santiago Peter Y. Lu Abstract Chaos arises in many complex dynamical systems, from weather to power grids, but is difficult to accurately model with data-driven methods such as machine learning emulators. While emulators are promising tools for accelerating simulations and solving inverse problems, they still struggle to learn chaotic dynamics, where sensitivity to initial conditions renders exact long-term forecasts infeasible, especially given noisy data. Recent work instead trains emulators to match the statistical properties of chaotic attractors, but these approaches often rely on handcrafted summary statistics or large, diverse multi-environment datasets. In this work, we propose a family of adversarial optimal transport objectives that can jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory. We theoretically analyze and experimentally validate a Sinkhorn divergence formulation (2-Wasserstein) and a WGAN-style dual formulation (1-Wasserstein) of our approach. Numerical experiments across a variety of chaotic systems, including ones with high-dimensional spatiotemporal chaos, show that emulators trained using our proposed objectives have significantly improved long-term statistical fidelity. 1 Introduction Chaos is a generic feature of high-dimensional nonlinear dynamical systems (Medio & Lines, 2001) and is fundamental to our understanding of statistical physics (Dorfman, 1999).

--- Page 2 ---

Learning to Emulate Chaos: Adversarial Optimal Transport Regularization learning to emulate chaos: adversarial optimal transport regularization. loss and an optimal transport cost that matches the summary statistic distribution of the model to the data, while the summary map is trained to maximize the same optimal transport cost. Unlike choosing a set of handcrafted summary statistics that may not be informative enough to constrain the model to a high-dimensional chaotic attractor, this adversarial objective for the summary map ensures it learns an optimally discriminative and therefore highly informative set of summary statistics. This results in an efficient method for ensuring the statistics of the trajectories produced by the emulator match the statistics of the true chaotic attractor. 1.1 Contributions 1 New statistics-based losses for emulating chaos. We introduce a family of adversarial optimal transport objectives that regularize the standard squared error prediction loss for emulator training. These new losses learn to emulate chaotic systems by simultaneously learning (i) an adversarial set of informative summary statistics and (ii) an emulator trained to match the distribution of these statistics to the noisy trajectory data. In practice, we use computationally tractable relaxations of the p-Wasserstein cost in our proposed loss, including a Wasserstein GAN-style loss for p = 1 and a Sinkhorn loss that provides an entropy-regularized relaxation valid for any choice of p ≥ 1.

--- Page 3 ---

Learning to Emulate Chaos: Adversarial Optimal Transport Regularization 2. Background and Problem Setting Notations Let (X, dX) and (Y, dY) be metric spaces. We denote by M+(X) the set of positive Radon probability measures on X. Given a measurable map T: X → Y and a measure µ ∈ M+(X), the pushforward measure T'µ ∈ M+(Y) is defined by T'µ(B):= µ(T−1(B)) for any measurable set B ⊆ Y. The set of joint probability measures on X × Y with marginals µ ∈ M+(X) and ν ∈ M+(Y) is defined by Π(µ, ν) = set of joint probability measures on X × Y with marginals µ ∈ M+(X) and ν ∈ M+(Y). The Kullback-Leibler divergence between two measures µ and ν is KL(µν). We denote by C(X,Y) the class of continuous maps X → Y. The∥ ·∥F is the Frobenius norm.

--- Page 4 ---

Learning to Emulate Chaos: Adversarial Optimal Transport Regularization 2.1 Dynamical Systems Setup Let (U, dU) be a compact metric space (e.g., a compact attractor). The true one-step dynamics is a Borelmeasurable map Φ: U → U admitting an invariant and ergodic probability measure µ, i.e. Φ'µ = µ.

Improvements for AI systems

  1. The improved AI system can learn high-quality summary statistics from a single noisy trajectory by jointly learning an adversarial set of informative summary statistics and (ii) an emulator trained to match the distribution of these statistics to the noisy trajectory data. This allows for emulating complex chaotic systems using only a single observation, which is currently difficult due to the sensitivity to initial conditions that causes standard MSE loss to become an increasingly poor objective for training emulators on chaotic dynamics over long rollouts.

  2. The system will achieve significantly improved long-term statistical fidelity by enforcing the distribution of learned statistics via an optimal transport cost, which is shown to be a robust objective against noise, as the analysis shows our statistics-based loss is much more robust to noise compared to standard mean squared error (MSE).

  3. The AI can utilize a horizon-dependent training strategy by applying horizondependent weighting: L(g, f):= X K k=1 wk L(k) MSE + λ Wd p S(f'µ, f'µˆ), where setting wk → 0 for k ≫ τmix down-weights noisedominated MSE terms. This ensures long-term statistical consistency even when trajectory level prediction is fundamentally impossible due to chaos, by down-weighting the MSE term after the mixing time.

  4. The system can accurately estimate chaotic dynamical properties, specifically the leading Lyapunov exponent (LLE), as Beyond attractor statistics, our approach also improves the dynamical properties of the emulator, such as the leading Lyapunov exponent (LLE), which represents the rate at which chaos becomes unpredictable. This allows for a direct comparison of how faithfully an emulator reproduces chaotic regimes by estimating LLE using the Benettin method.

  5. The system can be trained with explicit control over its learned complexity by implementing joint hinge regularization that enforces Lmin ≤ Lip(f) ≤ Lmax during training, penalizing when the summary map becomes geometrically irregular and preventing feature collapse, thus ensuring the learned geometry is stable.

Abstract

Chaos arises in many complex dynamical systems, from weather to power grids, but is difficult to accurately model with data-driven methods such as machine learning emulators. While emulators are promising tools for accelerating simulations and solving inverse problems, they still struggle to learn chaotic dynamics, where sensitivity to initial conditions renders exact long-term forecasts infeasible, especially given noisy data. Recent work instead trains emulators to match the statistical properties of chaotic attractors, but these approaches often rely on handcrafted summary statistics or large, diverse multi-environment datasets. In this work, we propose a family of adversarial optimal transport objectives that can jointly learn high-quality summary statistics and a physically consistent emulator from a single noisy trajectory. We theoretically analyze and experimentally validate a Sinkhorn divergence formulation (2-Wasserstein) and a WGAN-style dual formulation (1-Wasserstein) of our approach. Numerical experiments across a variety of chaotic systems, including ones with high-dimensional spatiotemporal chaos, show that emulators trained using our proposed objectives have significantly improved long-term statistical fidelity.

Sources

Related papers