Mamba Integrated with Physics Principles Masters Long-term Chaotic System Forecasting

arXiv:2505.23863 · cs.LG, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "PhyxMamba: Chaotic System Reconstruction from Short Context Observations with Generative State-Space Models".

Jane: The paper was written by Chang Liu, Bohao Zhao, Jingtao Ding, Huandong Wang and Yong Li from Tsinghua University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody! Today we’ve got a paper that sounds like it’s straight out of a sci-fi novel: “Mamba Integrated with Physics Principles Masters Long-Term Chaotic System Forecasting.” Jane, I’ve got to say, just the title alone has me hooked.

Jane: Same here, Tom! And honestly, the title is doing a lot of heavy lifting. “Chaotic systems” are all around us—weather, brain activity, even the stock market. The idea that we can forecast them long-term is a huge deal.

Tom: Right, and the team behind this is from Tsinghua University. We’ve got Chang Liu, Bohao Zhao, Jingtao Ding, Huandong Wang, and Yong Li. These folks are clearly deep into the machine learning and dynamical systems world.

Jane: And they’re not just throwing a neural network at the problem. They’re combining a modern AI model called Mamba with actual physics principles. That’s the “PhyxMamba” part.

Tom: So for our listeners who might not be familiar, Mamba is this new type of model—it’s a state-space model, which basically means it’s really good at handling sequences of data over time, like a time series. And it’s faster than the usual Transformer models we hear so much about.

Jane: Exactly. But here’s the kicker: chaotic systems are notoriously sensitive to initial conditions. A tiny error at the start can blow up exponentially. That’s the famous butterfly effect. So how do you forecast something like that long-term?

Tom: That’s the million-dollar question, and this paper tackles it head-on. They’re not just predicting the next few steps; they’re trying to capture the whole underlying structure of the system—the so-called “strange attractor.”

Jane: And that’s what makes this so exciting. It’s not just about getting the numbers right for a few time steps. It’s about understanding the shape and the rules of the chaos itself.

Tom: So, Jane, you’re saying this is more than just a better weather app? This could change how we model complex systems in science and engineering?

Jane: Absolutely. Think about climate models, predicting heart arrhythmias, or even understanding brain signals. If we can grasp the long-term behavior of chaos, we can make better decisions in all those fields.

Tom: I love it. And I know our listeners are going to want to hear the details of how they actually pulled this off. That’s coming up next.

Paper discussion segment 2: Jane: So, Tom, we’ve set the stage with the big picture. Now let’s get into the meat of the paper. How does PhyxMamba actually work?

Tom: Well, the first trick is something called time-delay embedding. Basically, they take a single variable from the system—like just the temperature reading—and they reconstruct the whole multi-dimensional state of the system from that one line of data.

Jane: It’s like if you only saw the shadow of a dancer on a wall, but you could still figure out the entire choreography from the way the shadow moves. The paper uses Takens’ theorem to do this mathematically.

Tom: Exactly. So they’re not just feeding the raw numbers into the model. They’re building a richer, physics-informed representation of the system’s attractor. That’s the “physics principles” part of the title.

Jane: Then they feed that into Mamba, but they don’t just train it to predict the next time step. They train it to generate the next “patch” of data, kind of like how a language model predicts the next word.

Tom: And they go even further. They also train it to predict multiple patches ahead at once. That forces the model to learn the global dynamics, not just the short-term wiggle.

Jane: Right, and this is where it gets clever. They call it “student forcing.” Instead of always giving the model the correct answer during training, they let it make mistakes and then learn to correct itself. This is crucial for long-term forecasting because errors can accumulate.

Tom: And to make sure the model doesn’t just drift off into some made-up fantasy, they add a regularizer based on Maximum Mean Discrepancy. That’s a fancy way of saying they check if the distribution of the model’s predictions matches the distribution of the real system.

Jane: So it’s not enough for the model to be close on average. It has to produce states that look like they could actually come from the real chaotic system. The shape of the attractor has to be right.

Tom: That’s a really elegant way to keep the model honest. And the results? They’re pretty stunning. On the Rossler system, they get a valid prediction time of nearly ten Lyapunov times. That’s ten times the natural timescale of chaos.

Jane: And on real-world EEG data, they’re getting similar results. That’s brain activity, which is incredibly noisy and complex. The fact that they can capture its long-term structure is a big deal.

Tom: So, Jane, it sounds like they’ve cracked the code on combining physics with modern AI. But what does this mean for the people actually trying to use these models in the real world? That’s a question for our next segment.

Paper discussion segment 3: Tom: We’re back, and we’ve got our senior researcher, Lu, and our engineer, Meng, with us to dig into the practical side of PhyxMamba. Lu, what’s the big deal here from a research perspective?

Lu: Thanks, Tom. The big deal is that they’re solving a problem that’s been a bottleneck for a long time. Most models need a ton of data that covers the full range of the system’s behavior. But this paper shows you can train on just one Lyapunov time of data—that’s a very short window—and still forecast for ten times that long.

Meng: And that’s not just a theoretical win. From an engineering standpoint, that’s huge. It means we can deploy these models in situations where we don’t have years of historical data. Think about a new sensor in the ocean or a medical monitor that’s just been installed.

Jane: So, Meng, you’re saying the data efficiency is a game-changer for real-world deployment?

Meng: Absolutely. And it’s not just about data. The model is also computationally efficient. Mamba is linear-time, which means it scales much better than the quadratic cost of Transformers. So you can run longer forecasts without blowing up your compute budget.

Lu: And that efficiency lets them do something really clever: they use a residual stacking architecture. Each layer of Mamba learns to model a different component of the dynamics. It’s like decomposing the system’s motion into distinct parts, which makes the model more interpretable.

Tom: Interpretable? That’s a word we don’t hear enough in deep learning. So you can actually see what the model is learning?

Lu: To a degree, yes. Instead of a black box, you have a model that’s explicitly breaking down the dynamics into hierarchical pieces. That aligns with how we think about complex systems in physics.

Meng: And from my side, the robustness results are what really catch my eye. They tested it with added noise and with less training data, and the long-term statistics—the shape of the attractor—stayed stable. That’s the sign of a model that’s learning the actual physics, not just memorizing the training set.

Jane: So it’s not just a better predictor; it’s a more trustworthy one. That’s what we need for high-stakes applications.

Tom: And that’s what makes this paper so impactful. It’s not just a step forward in AI; it’s a bridge between AI and the fundamental laws of nature. We’ll wrap this up with our final thoughts in just a moment.

Conclusion: Tom: Alright, we’re wrapping up our discussion on “Mamba Integrated with Physics Principles Masters Long-Term Chaotic System Forecasting.” Jane, what’s the one thing you want our listeners to remember?

Jane: I think it’s that this paper shows we can forecast chaos, not just react to it. By combining Mamba’s efficiency with physics-based representations and smart training strategies, they’ve built a model that understands the underlying structure of chaotic systems.

Tom: And that’s a big deal for the future. We’re talking about better climate models, more reliable medical monitors, and maybe even a deeper understanding of complex financial systems.

Lu: And the implications go beyond just prediction. This approach of embedding physical principles into generative models could be a blueprint for other scientific fields. It’s a way to make AI not just powerful, but also grounded in reality.

Meng: And from a practical standpoint, it’s efficient and data-hungry in a good way. That means we can start using this in real-world systems much sooner than we might have thought.

Tom: So, we’ve got a paper that’s both scientifically deep and practically useful. That’s a rare combination. We’ll be keeping an eye on how this line of research develops.

Jane: Absolutely. And with that, we’re saying goodbye to this paper. Thanks for joining us, and we’ll be back soon to break down the next big idea from the arXiv.

Chang Liu, Bohao Zhao, Jingtao Ding, Huandong Wang, Yong Li

Tsinghua University

cs.LG, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 58/100

The gist: Long-term forecasting of chaotic systems remains a fundamental challenge due to the intrinsic sensitivity to initial conditions and the complex geometry of strange attractors.

Key concepts

Mamba
Mamba is a new type of state-space model used in AI. It is highly effective at processing sequences of data over time (like time series) and is faster than traditional Transformer models.
Chaotic Systems
These systems, such as weather or brain activity, are extremely sensitive to initial conditions. A tiny error can cause the prediction to grow exponentially, making long-term forecasting a major challenge.
Time-delay embedding
This technique allows researchers to reconstruct a complex system's multi-dimensional state from just one variable's data. It helps capture the entire structure of the system using mathematical principles.

Terminology

Summary

Long-term forecasting of chaotic systems remains a fundamental challenge due to the intrinsic sensitivity to initial conditions and the complex geometry of strange attractors. The paper states: "Conventional approaches, such as reservoir computing, typically require training data that incorporates long-term continuous dynamical behavior to comprehensively capture system dynamics. While advanced deep sequence models can capture transient dynamics within the training data, they often struggle to maintain predictive stability and dynamical coherence over extended horizons."

The paper addresses the specific problem of forecasting the long-term dynamics of chaotic systems given short-term historical observations on their states, which the authors note remains a largely underexplored scientific problem in existing research. Three fundamental challenges are identified: (i) Extracting complex dynamical information from short observation windows, (ii) Maintaining long-term predictive stability and dynamical coherence, and (iii) Reproducing key statistical properties and attractor geometry.

The authors propose PhyxMamba, a framework that integrates a Mamba-based state-space model with physics-informed principles to forecast long-term behavior of chaotic systems given short-term historical observations on their state evolution. The framework consists of three main components:

The paper states: Rather than relying on neural network-based autoencoders for representation learning, we propose a physics-informed pipeline to construct high-dimensional representations for chaotic systems. Using Takens' embedding theorem, the authors construct high-dimensional representations through time-delay embedding: given a univariate trajectory x ∈ RT from a chaotic system's attractor with T observed steps, we define two hyperparameters: the embedding dimension m and the time delay τ, producing representations zt = (xt−(m−1)τ, xt−(m−2)τ, · · ·, xt) ∈ Rm. The parameters m and τ are chosen using CC methods. The combined representations from different variables are defined as Z ∈ RV ×T ×m. The representations are then partitioned into patches and mapped to latent embeddings via linear mapping.

The paper states: we employ a generative training strategy, enabling the model to predict the next patch autoregressively. The prediction for patch Pi is formulated as P̂i = f (P0, P1, · · ·, Pi−1). The authors adopt a residual stacking architecture designed to decompose the underlying system dynamics into distinct components, with each Mamba layer learning to model different aspects of the dynamics. For the l-th Mamba layer: hi(l) = Mamba(l) (r0:i(l−1)), Êi(l) = Decoder(l) (hi(l)), ri(l) = ri(l−1) − Êi(l).

The paper also introduces a multi-patch prediction (MPP) objective, encouraging the model to capture the global dynamics of the underlying system. This uses M auxiliary modules to predict M subsequent patches, with each module containing a dedicated Mamba layer, and a fusion projection layer ψ(·). The training objective combines next-patch prediction loss and multi-patch prediction losses.

The paper states: To mitigate this issue, we introduce an additional student-forcing training phase, wherein the model autoregressively generates W patches based on its historical predictions. The authors employ Maximum Mean Discrepancy (MMD) to quantify the distance between empirical distributions phist, ppred, and pgt formed by these trajectories. The regularization term is defined as: Lreg = MMD2(phist, ppred) + λc MMD2(pgt, ppred), where the first term enforces statistical consistency between the historical dynamics and the model's predictions, while the second aligns the predicted state distribution with that of the ground truth. The kernel κ is implemented as a mixture of rational quadratic kernels.

Datasets: The method is evaluated on three standard simulated chaotic systems: the 3-dimensional Lorenz63, Rössler, and 5-dimensional Lorenz96, plus real-world datasets: the 5-dimensional Electrocardiogram (ECG) and 64-dimensional Electroencephalogram (EEG). Data is resampled according to their Lyapunov time (TL), where each Lyapunov time corresponds to 30 time steps for most systems, and 60 time steps for EEG.

Evaluation Metrics: The paper uses 1-step Error, symmetric mean absolute percentage error (sMAPE), and valid prediction time (VPT) for point-wise accuracy, and the correlation dimension error (Dfrac) and KL divergence between attractors (Dstsp) for statistical fidelity.

Baselines: Two categories are compared: physics-informed dynamical systems models including Koopa, PLRNN, nVAR, PRC, and HoGRC, and general time series forecasting models including NBEATS, NHiTS, CrossFormer, PatchTST, TimesNet, TiDE, iTransformer, DLinear, NSFormer, FEDFormer, and AutoFormer. Zero-shot performance of Timer and Chronos is also investigated.

Training Setup: During training, our model utilizes data spanning one Lyapunov time (1TL) to perform next-patch prediction with teacher forcing. The model then autoregressively predicts the subsequent TL trajectory based on the preceding TL segment in the student forcing stage. For evaluation, models are given observation trajectories of length TL and tasking them with predicting the subsequent trajectories spanning 10TL.

The paper reports: Our model consistently achieves state-of-the-art point-wise forecasting accuracy across all evaluated datasets. Specifically, our model attains a VPT of 9.83TL and 9.43TL on the Rossler simulation and the real-world EEG dataset, respectively, significantly surpassing baselines. The model also excels at preserving the long-term statistical characteristics and geometric integrity of the strange attractors, with Dfrac of 0.049 and Dstsp of 0.034 on the Rossler dataset versus best baseline values of 0.069 and 0.137.

The paper notes: "Strong short-term predictive accuracy does not necessarily guarantee robust long-term forecasting performance. Models exhibiting lower 1-step prediction errors, such as HoGRC, PRC, and NHiTS, often show rapid degradation in their forecasts."

The ablation study reveals: removing physics-informed representation results in the most substantial degradation of the attractor's long-term geometric integrity, while ablation of student-forcing led to the most acute decline in point-wise accuracy. The paper concludes: all architectural designs and training strategies are crucial for achieving the superior performance demonstrated by our proposed method.

Against Noise: "although point-wise accuracy (sMAPE@1) deteriorates notably when σ > 0.1, the long-term consistency metric (Dstsp) remains remarkably stable, demonstrating the model's ability to faithfully preserve the geometric structure and essential statistical invariants of the system's strange attractor."

Against Training Data Ratio: "our approach exhibits significantly higher data efficiency than the best baseline methods. For instance, on the Rossler system, our model requires only 20% of the full training dataset to outperform the best baseline in both point-wise accuracy and long-term geometric consistency."

The paper concludes: "By synergistically integrating a Mamba-based state-space model with physics-informed embedding, our approach effectively achieves point-wise state forecasting accuracy and reproduces key statistical properties of systems. Through extensive experiments on both simulation and real-world chaotic system datasets, we demonstrate the effectiveness in chaotic system forecasting of our model design and superior performance compared with baseline models. The work paves the way for future research in modeling complex dynamical systems under observation-scarce conditions, with potential applications in climate science, neuroscience, epidemiology, and beyond."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved systems can do:

Improvement 1: Physics-Informed State Representation for Time Series Models

  • Implementation: Add a pre-processing layer that applies Takens’ time-delay embedding to each input variable before feeding it into the model. This creates a higher-dimensional representation (Z ∈ R V×T×m) that captures the underlying attractor manifold. Use the CC method to automatically select embedding dimension (m) and delay (τ).

  • What the improved AI system can do: It can learn the global dynamics of a chaotic system from a very short observation window (e.g., 1 Lyapunov time) instead of requiring long, continuous trajectories. This makes it effective for real-world scenarios where only brief historical data is available (e.g., early warning systems for equipment failure, short-term climate anomaly detection).

Improvement 2: Generative Next-Patch Training with Multi-Patch Prediction (MPP)

  • Implementation: Replace standard sequence-to-sequence training with a decoder-only, autoregressive next-patch prediction objective. Add M auxiliary prediction heads that forecast M future patches simultaneously during training. The loss function combines next-patch loss (L next) and multi-patch losses (L MPP) with a weighting factor λp.

  • What the improved AI system can do: It learns the underlying generative process rather than just a direct input-output mapping. This enables the model to maintain predictive stability and dynamical coherence over long horizons (e.g., 10 Lyapunov times), preventing the error accumulation that plagues standard autoregressive models. It also improves data efficiency, requiring only 20–40% of training data to outperform baselines trained on full datasets.

Improvement 3: Two-Stage Training with Student Forcing and Distribution Matching

  • Implementation: After the initial teacher-forcing phase, add a second training stage where the model autoregressively generates W patches based on its own predictions (student forcing). During this stage, add a regularization term (L reg) that uses Maximum Mean Discrepancy (MMD) with a mixture of rational quadratic kernels to minimize the distance between the predicted state distribution and both the historical and ground-truth distributions.

  • What the improved AI system can do: It becomes robust to its own prediction errors during inference, avoiding the train-test mismatch that causes performance degradation. The MMD regularization ensures the model reproduces the invariant measure and fractal geometry of the strange attractor, preventing collapse into trivial states or spurious attractors. This is critical for applications requiring statistical fidelity, such as generating realistic synthetic data for stress-testing financial models or simulating disease spread for public health planning.

Improvement 4: Residual Stacking Architecture for Decomposed Dynamics

  • Implementation: Use a stack of Mamba layers where each layer l learns the residual between the input and the cumulative prediction of previous layers. Each layer outputs a component (Ê i(l)) that represents a distinct aspect of the dynamics, and the final prediction is the sum of all components.

  • What the improved AI system can do: It can decompose complex, composite system dynamics into interpretable components, each modeled by a separate layer. This improves forecasting accuracy by allowing the model to capture both fast and slow timescales separately, and it provides a level of interpretability that is useful for scientific discovery (e.g., identifying which components drive short-term vs. long-term behavior).

Improvement 5: Adaptive Observation Length at Inference

  • Implementation: Design the model to accept variable-length input sequences at inference time (as long as the length is a multiple of the patch size). This is enabled by the generative next-patch training paradigm.

  • What the improved AI system can do: It can generate accurate long-term forecasts even when the available historical context is shorter than the training window. The paper shows that the model maintains stable long-term geometric fidelity (Dstsp) even with only 0.2TL of input, and in some cases, shorter inputs yield better short-term accuracy than longer ones. This flexibility is crucial for real-time applications where data may be incomplete or delayed.

Summary of Capabilities of the Improved AI System:

  • Long-horizon forecasting: Predicts chaotic system behavior for up to 10 Lyapunov times with high point-wise accuracy (e.g., VPT of 9.83TL on Rossler, 9.43TL on EEG).

  • Statistical fidelity: Reproduces the attractor’s correlation dimension and KL divergence with errors as low as 0.049 and 0.034, respectively, on the Rossler system.

  • Data efficiency: Achieves superior performance with only 20–40% of the training data required by state-of-the-art baselines.

  • Noise robustness: Maintains stable long-term statistical properties even under significant observational noise (σ up to 0.5).

  • Flexibility: Can operate with variable-length historical inputs and is computationally efficient (e.g., 3.65M parameters, 184ms inference time for 10TL forecasting on Lorenz63).

Abstract

Long-term forecasting of chaotic systems remains a fundamental challenge due to the intrinsic sensitivity to initial conditions and the complex geometry of strange attractors. Conventional approaches, such as reservoir computing, typically require training data that incorporates long-term continuous dynamical behavior to comprehensively capture system dynamics. While advanced deep sequence models can capture transient dynamics within the training data, they often struggle to maintain predictive stability and dynamical coherence over extended horizons. Here, we propose PhyxMamba, a framework that integrates a Mamba-based state-space model with physics-informed principles to forecast long-term behavior of chaotic systems given short-term historical observations on their state evolution. We first reconstruct the attractor manifold with time-delay embeddings to extract global dynamical features. After that, we introduce a generative training scheme that enables Mamba to replicate the physical process. It is further augmented by multi-patch prediction and attractor geometry regularization for physical constraints, enhancing predictive accuracy and preserving key statistical properties of systems. Extensive experiments on simulated and real-world chaotic systems demonstrate that PhyxMamba delivers superior forecasting accuracy and faithfully captures essential statistics from short-term historical observations.

Sources

Related papers