ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks

arXiv:2502.00803 · cs.LG · Submitted 2026-08-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks".

Jane: The paper was written by Yuezhou Ma, Haixu Wu, Hang Zhou, Huikun Weng, Jianmin Wang et al. from School of Software, BNRist, Tsinghua University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the physics and machine learning world, and it's called "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks." Jane, I have to say, just the title alone got me excited.

Jane: Oh, absolutely, Tom. And for our listeners who might not be deep in the weeds of this, let's break that title down. Physics-informed neural networks, or PINNs, are these clever AI models that try to solve physics equations by learning from the equations themselves, not just from data. But they have this frustrating habit of failing on some pretty simple problems.

Tom: And that's where "propagation failures" comes in. The paper is saying that the correct answer, the supervision, gets stuck at the boundaries and never really spreads into the middle of the problem space. It's like the model learns the edges of the puzzle but never figures out the center.

Jane: Exactly. And the authors, they're from Tsinghua University, which is a powerhouse for this kind of research. They're not just describing the problem, they're claiming to have figured out why it happens and how to fix it. That's a big deal.

Tom: A huge deal. Because if you can solve these simple failures, you can start tackling much bigger, more complex physics problems. We're talking about simulating fluid dynamics, weather patterns, maybe even designing new materials. The implications are massive.

Jane: Right, and the fact that they're calling it "demystifying" suggests they've found the root cause, not just a patch. That's what I'm most curious about. How did they figure out what's actually going wrong inside the model?

Tom: Well, we're going to get into exactly that in the next segment. They have a pretty wild theory about gradients and correlations, and it's going to change how you think about these networks.

Jane: I can't wait. Stick around, because we're just getting started with "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks."

Summary: Tom: So we've set the stage. The paper is "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks," and Jane, you were asking about the root cause. Let's get into the summary of what these researchers actually did.

Jane: Right. So, the key insight here is that they compared PINNs to the old-school way of solving these equations, the finite element method. In the old method, you break the space into a connected mesh, like a fishing net, and each point is directly linked to its neighbors. When you update one point, it pulls on the points around it.

Tom: Like a physical net, right? Pull one knot and the whole net shifts.

Jane: Precisely. But a PINN, the way it's usually built, treats every point in the space as an island. It processes each coordinate independently. There's no direct link between neighboring points. The paper proves that this lack of connection is the core problem.

Tom: And they didn't just guess this. They actually defined a mathematical way to measure this "propagation." They call it the stiffness coefficient, borrowing from physics, and they show it's directly equal to something called the gradient correlation between two nearby points.

Jane: And that gradient correlation, in plain English, is a measure of whether updating the model to fix the answer at one point also helps fix the answer at a nearby point. If the gradients are pointing in totally different directions, the model is learning contradictory things for two points that should be almost the same.

Tom: So, the failure isn't just about the loss function being hard. It's about the architecture of the model itself preventing information from flowing. That's a fundamental shift in how we understand the problem.

Jane: It really is. And this gives them a precise tool. Instead of just looking at a vague error map, they can now look at the gradient correlation map and see exactly where the model is going to fail, before it even fails.

Tom: That's the diagnostic power. But the real question is, what do you do about it? And that's what we're going to talk about next, because they didn't just stop at the diagnosis. They built a whole new architecture to fix it.

Jane: And that architecture is the "Pro" in ProPINN. Let's get into it.

Improvements: Tom: Welcome back. We're deep in "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks," and we've established the problem: the model's architecture isolates each point. So, Jane, how do these authors fix that?

Jane: They introduce something called "multi-region mixing." The idea is to force the model to look at a point and its immediate neighborhood all at once. Instead of just feeding the model a single coordinate, they perturb that coordinate, creating a small cloud of points around it.

Tom: So, for every point you want to solve, you also create a few nearby points, and you process them together.

Jane: Exactly. And crucially, they do this in a "differential" way. The gradients from those nearby points are allowed to flow back and update the model together. This is the key move. By uniting the gradients of these region points, they artificially boost the gradient correlation that we talked about earlier.

Tom: So they're not just adding more data points; they're creating a structure where the model is forced to learn a consistent answer for a whole area, not just isolated spots.

Jane: Right. And they prove mathematically that this design improves the gradient correlation between nearby points. It's a theoretical guarantee, not just a trick that seems to work.

Tom: And the efficiency is a big part of this too. They kept the extra computation minimal by only doing this region mixing in a small, lightweight part of the network. They didn't build a huge, heavy model like some other approaches.

Jane: That's a great point, Tom. Because the previous state-of-the-art models that tried to capture these correlations used Transformers, which are notoriously slow and memory-hungry. ProPINN gets the benefit of that correlation without the massive computational cost.

Tom: So, they've got the theory, they've got the architecture, and they've got the efficiency. The real test is whether it actually works on real problems. And that's what we're going to look at in the next segment, the actual results from the first page of the paper.

Jane: The numbers are pretty convincing, so stay tuned.

First Page: Tom: We're back with "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks," and we've talked about the theory and the design. Now, let's look at the actual results, the proof in the pudding. Jane, what did they find?

Jane: They tested it on a bunch of standard physics problems, the ones that usually break vanilla PINNs. And the improvement is stark. On the Convection problem, which is a classic failure case, the vanilla PINN has a relative error of about zero point seven seven eight. ProPINN brings that down to zero point zero one eight.

Tom: That's not an incremental improvement. That's a forty-fold reduction in error. It's like going from a blurry guess to a sharp photograph.

Jane: And it's not just that one problem. Across all four standard benchmarks, they consistently beat the second-best model by a huge margin. On the Allen-Cahn equation, they got a sixty-three percent relative improvement over the previous best. On the 1D-Wave equation, it was a sixty-nine percent improvement.

Tom: And these aren't easy problems. These are the ones that have been used to demonstrate how PINNs can fail. So, they've essentially solved the failure modes that have been plaguing the field.

Jane: But they didn't stop there. They also tested it on much harder, real-world problems like the Navier-Stokes equations for fluid dynamics. This is simulating things like the Karman vortex street, the swirling patterns you see behind a cylinder in a flow.

Tom: And even there, they're getting a forty-four percent improvement over the previous best model on the Karman Vortex task. That's huge for practical applications.

Jane: And here's the kicker, Tom. They did all this while being two to three times faster than the Transformer-based models. So, they're not just more accurate; they're more efficient too. That's a rare combination.

Tom: So, the first page alone is packed with evidence that this architecture is a game-changer. It's faster, more accurate, and it solves the fundamental problem they identified. I'm really excited to see where this goes.

Jane: Me too. And we're going to wrap up our thoughts on the whole paper in our final segment.

Conclusion: Tom: And that brings us to the end of our discussion on "ProPINN: Demystifying Propagation Failures in Physics-Informed Neural Networks." Jane, it's been a fantastic paper to break down.

Jane: It really has, Tom. We started with a mystery: why do these clever AI models fail on simple physics? And we ended with a clear, mathematical answer. The problem was the architecture, and the solution is a new way of building the network that forces it to learn connected, consistent solutions.

Tom: And the implications go way beyond just fixing a bug. If we can reliably solve these equations, we can build better simulations for weather forecasting, for designing aircraft, for understanding blood flow in the human body. The potential for real-world impact is enormous.

Jane: Absolutely. And the fact that they did it with a lightweight, efficient design means it's not just a lab curiosity. It's something that can be deployed in real engineering workflows. That's what makes this paper so significant.

Tom: So, we say goodbye to ProPINN, but we're definitely not saying goodbye to the ideas it introduced. Gradient correlation is going to be a concept we hear a lot more about in the future.

Jane: For sure. It gives us a new lens to look at all kinds of neural network training problems, not just physics. Thanks for joining us, everyone.

Tom: And stay tuned, because we've got another exciting paper coming up next. We'll see you then.

Yuezhou Ma, Haixu Wu, Hang Zhou, Huikun Weng, Jianmin Wang, Mingsheng Long

School of Software, BNRist, Tsinghua University

cs.LG

Submitted: 2026-08-09

Code: https://github.com/PredictiveIntelligenceLab/jaxpi

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 70/100

Key concepts

Physics-Informed Neural Networks (PINNs)
AI models designed to solve physics equations by learning directly from the governing physical equations, rather than relying solely on measured data. They are used for simulating complex systems like fluid dynamics and weather patterns.
Propagation Failures
A common issue where PINNs fail to accurately solve problems because the correct information or supervision gets stuck at the boundaries and does not spread into the center of the problem space.
Gradient Correlation
A mathematical measure used to determine if updating a model's answer at one point also helps fix its answer at a nearby point. Low correlation indicates that the model is learning contradictory information.
Multi-region Mixing
The core architectural improvement in ProPINN. It forces the model to process not just single coordinates, but small clouds of nearby points together, allowing gradients to flow back and update the model consistently across a region.

Terminology

Summary

Summary

This paper provides a formal and in-depth study of the propagation failure phenomenon in Physics-Informed Neural Networks (PINNs), which is a key optimization challenge where correct supervision from initial or boundary conditions fails to reach interior domain points. The authors go beyond intuitive loss-map-based explanations and establish a theoretical root cause for this failure, leading to a new architecture called ProPINN.

Problem and Motivation: PINNs solve PDEs by embedding equation constraints, initial conditions, and boundary conditions into a loss function. However, they often fail on simple PDEs, a phenomenon termed PINN failure modes. Previous work by Daw et al. described propagation failure intuitively as the failure of correct supervision to propagate from initial states or boundaries to the interior domain. The authors note that this understanding is insufficient: a formal and rigorous definition of 'propagation' processes during PINN optimization and the root cause of propagation failure are still underexplored. They also point out that loss-oriented methods, such as sampling strategies, may not solve the propagation issue at its root, as the residual loss distribution can be too dispersed to identify the failure's origin.

Theoretical Contribution: The paper draws inspiration from classical numerical methods, specifically Finite Element Methods (FEMs), which do not suffer from propagation issues. The key insight is that FEMs discretize the domain into connected meshes where nearby points directly interact, whereas PINNs usually treat the input domain as a set of independently processed collocation points, reducing interaction between outputs on nearby positions.

The authors formalize this by defining a stiffness coefficient for PINNs, analogous to the stiffness matrix in FEMs. This leads to their central theorem:

Theorem 3.6 (Gradient-correlation): For a PINN u theta and adjacent points x, x' in, the PINN stiffness coefficient is equivalent to the gradient correlation between the two points:

[

D PINN(x, x') = G u theta(x, x') = d u theta over d theta x, d u theta over d theta x'.

]

This proves that weak propagation between adjacent points... is equivalently characterized by a small gradient correlation. The authors explain why this failure occurs: since the gradient tensors are high-dimensional (parameter size usually > 10 cubed), they are easy to be orthogonal in the high-dimensional space, especially when different positions are independently optimized. This also provides insight into why PINNs cannot benefit from large models, as a larger parameter size increases the likelihood of orthogonal gradients.

ProPINN Architecture: Based on this theoretical finding, the authors propose ProPINN, a new architecture designed to enhance gradient correlation by uniting region gradients. Its key component is a multi-region mixing mechanism with two main parts:

  1. Differential Perturbation: Given a collocation point x, it is augmented by perturbing its position within multiscale regions. The augmented points are processed by a shared lightweight projection layer P. This design is differential (gradients flow through it), which is crucial: from the backward perspective... the differential perturbation design can also aggregate gradients of collocation points within multiscale regions, thereby explicitly enhancing their gradient correlation.

  2. Multi-Region Mixing: Instead of using computationally expensive attention mechanisms like Transformer-based models, ProPINN averages representations within each region to create region representations, then mixes them with a simple MLP layer M. The authors argue that the relation among different coefficient PDEs is much more steady than complex point-to-point dependencies, allowing for a simpler and more efficient design.

The paper proves that this design improves gradient correlation under a reasonable assumption of positive local correlation (Theorem 3.10).

Experimental Results: ProPINN is evaluated on six PDE-solving tasks, including standard benchmarks (Convection, 1D-Reaction, Allen-Cahn, 1D-Wave) and complex fluid dynamics (Karman Vortex, Fluid Dynamics).

  • Standard Benchmarks: ProPINN achieves state-of-the-art performance, significantly outperforming all baselines including Transformer-based models (PINNsFormer, SetPINN). For example, on the Convection task, ProPINN achieves an rMAE of 0.018 compared to 0.023 for PINNsFormer, a 22% improvement. On Allen-Cahn, it achieves 0.036 rMAE versus 0.098 for PirateNet, a 63% improvement. The paper notes that only architectures that consider interactions among multiple points (PINNsFormer, SetPINN and ProPINN) consistently successfully converge, while single-point-processing models fail on Convection.

  • Complex Physics: On the challenging Navier-Stokes equations, ProPINN again achieves the best performance. For Karman Vortex, it achieves an rMAE of 0.161, a 44% improvement over SetPINN (0.287). For Fluid Dynamics, it achieves 0.1834 rMAE, a 22% improvement over FLS (0.2362). Notably, Transformer-based models (PINNsFormer, SetPINN) suffer from training instability (Nan) on this task, while ProPINN performs well.

  • Efficiency: ProPINN is about 2-3× faster than Transformer-based models (PINNsFormer, SetPINN) and comparable to single-point-processing PINNs (QRes, FLS) while providing over 60% error reduction.

  • Model Analysis: Ablations confirm that both multi-region mixing and differential perturbation are essential. The paper shows that the non-differential perturbation will seriously damage the model's performance, confirming that uniting gradients of region points is more important for PINN optimization. ProPINN also demonstrates favorable scalability, while vanilla PINN's performance drops when scaled up. Gradient correlation analysis shows that ProPINN maintains a high, stable gradient correlation during training, whereas vanilla PINN's correlation progressively worsens.

Conclusion: The paper concludes that ProPINN can reliably mitigate PINN failure modes and achieve state-of-the-art with 46% relative gain in typical PDE-solving tasks with favorable efficiency. It also presents better scalability and provides a precise, quantifiable criterion (gradient correlation) for identifying PINN failures.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems, along with what the improved system can do:


1. Architecture-Level Fix for Propagation Failure in PINNs

  • Improvement: Replace the standard single-point-processing MLP backbone in physics-informed neural networks with the ProPINN architecture, which includes a differential perturbation layer and a multi-region mixing mechanism.

  • What the improved system can do:

  • Solve PDEs that previously failed (e.g., Convection with β=50, 1D-Reaction, Allen-Cahn) with 46% relative error reduction compared to the best Transformer-based models.

  • Maintain high accuracy even when the PDE solution has steep gradients or rapid temporal changes, where vanilla PINNs collapse to trivial solutions.

2. Gradient-Correlation-Aware Training

  • Improvement: Introduce a new training criterion based on the paper’s Theorem 3.6: monitor and optimize the gradient correlation between nearby collocation points (not just residual loss). This can be used as a diagnostic tool or a regularizer.

  • What the improved system can do:

  • Detect propagation failure before it manifests in the error map, allowing early intervention.

  • Provide a quantitative measure to decide where to add collocation points or adjust loss weights, rather than relying on diffuse residual loss maps.

3. Efficient Multi-Scale Region Augmentation

  • Improvement: Augment each input point with multiscale perturbations (e.g., 3 scales with sizes 0.01, 0.05, 0.09) and share a lightweight projection layer across all augmented points, then pool and mix region representations.

  • What the improved system can do:

  • Achieve 2–3× faster training than Transformer-based PINNs (PINNsFormer, SetPINN) while maintaining or exceeding their accuracy.

  • Scale to larger models without the performance degradation seen in vanilla PINNs (which suffer from orthogonal gradients as parameter count grows).

4. Robustness to High-Order Derivatives

  • Improvement: Use ProPINN’s architecture for PDEs involving second-order or higher derivatives (e.g., 1D-Wave equation).

  • What the improved system can do:

  • Solve high-order PDEs with 71% relative error reduction compared to the second-best model, avoiding the optimization instability that plagues attention-based models under high-order loss landscapes.

5. Integration with Existing Training Strategies

  • Improvement: Combine ProPINN with complementary techniques: R3 sampling, loss reweighting, and RoPINN optimizer.

  • What the improved system can do:

  • Further boost performance on complex tasks (e.g., 1D-Reaction) without architectural changes, confirming orthogonality and composability.

Capability Before (Vanilla PINN) After (ProPINN-based)


Solve Convection (β=50) Fails (rMAE ≈ 0.78) Succeeds (rMAE ≈ 0.018)

Solve 1D-Wave (2nd-order) Fails (rMAE ≈ 0.33) Succeeds (rMAE ≈ 0.016)

Solve Karman Vortex (Navier-Stokes) Fails (rMAE ≈ 13.08) Succeeds (rMAE ≈ 0.161)

Training speed vs. Transformers — 2–3× faster

Model scalability Degrades with size Stable and improves

Failure detection Only after error appears Predictable via gradient correlation

If you want, I can also provide the exact architectural pseudocode or PyTorch/JAX implementation details for the ProPINN layer to integrate into your existing PINN codebase.

Abstract

Physics-informed neural networks (PINNs) have earned high expectations in solving partial differential equations (PDEs), but their optimization usually faces thorny challenges due to the unique derivative-dependent loss function. By analyzing the loss distribution, previous research observed the propagation failure phenomenon of PINNs, intuitively described as the correct supervision for model outputs cannot ''propagate'' from initial states or boundaries to the interior domain. Going beyond intuitive understanding, this paper provides a formal and in-depth study of propagation failure and its root cause. Based on a detailed comparison with classical finite element methods, we ascribe the failure to the conventional single-point-processing architecture of PINNs and further prove that propagation failure is essentially caused by the lower gradient correlation of PINN models on nearby collocation points. Compared to superficial loss maps, this new perspective provides a more precise quantitative criterion to identify where and why PINN fails. The theoretical finding also inspires us to present a new PINN architecture, named ProPINN, which can effectively unite the gradients of region points for better propagation. ProPINN can reliably resolve PINN failure modes and significantly surpass advanced Transformer-based models with 46% relative promotion.

Sources

Related papers