Learning in PINNs: Phase transition, diffusion equilibrium, and generalization

arXiv:2403.18494 · cs.LG · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Learning in PINNs: Phase transition, diffusion equilibrium, and generalization".

Jane: The paper was written by Sokratis J. Anagnostopoulosa, Juan Diego Toscanob, Nikolaos Stergiopulosa and George Em Karniadakis from EPFL and Brown University, Division of Applied Mathematics and Brown University, School of Engineering, Brown University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title and Authors: Tom: So, we've established that "Learning in PINNs: Phase transition, total diffusion, and generalization" is the name of the game here. It’s not just about solving a differential equation; it’s about *how* the network learns to solve it.

Jane: Think of a phase transition like water freezing—it’ goes from one state to another, and you can't just smoothly transition between phases for a certain property. The authors are showing that AI learning has these distinct "states."

Lu: And when they introduce the concept of "total diffusion," they're suggesting that the most stable, robust phase where everything aligns perfectly is not necessarily the start or end point, but a specific equilibrium state in between it is becoming critical.

Meng: It’s a really good framework because we usually just want things to converge as fast as possible, but this paper suggests convergence might be slow until that specific moment of total diffusion triggers a huge jump in quality.

Lalam: I find the idea of generalization being tied to this specific phase particularly profound, suggesting that how well our AI models understand fundamental physical laws is directly linked to their internal stability during training.

Tom: It seems like we’re looking at the authors' work not just as a mathematical exercise, but as a blueprint for how neural networks are fundamentally designed to operate.

Jane: Which leads us straight into understanding what these phases look like in the second segment of our discussion.

Summary and Core Findings: Tom: We’ve seen that this paper details three main learning phases: fitting, diffusion, and total diffusion. But how does it actually measure these phases?

Jane: The authors use a metric called the Signal-to-Noise Ratio or SNR to track it. Think of the signal as the clear direction of all the training data pointing toward a solution, and noise as everything else that pulls them off course.

Lu: The theory suggests that initially, we are in a "fitting" phase where things are highly structured and deterministic, but eventually, we drift into this noisy "diffusion" period.

Meng: The crucial discovery is what happens when the SNR goes through a sharp transition into total diffusion—that’s when the signal and noise both align perfectly for practical purposes, which means our training steps are consistent across all batches.

Lalam: It's a powerful idea that we can quantify this consistency, proving that the moment of maximum stability is also where the greatest potential for understanding physical reality lies.

Tom: And when this aligns with "gradient homogeneity," as they call it, it means every sample in the training set is agreeing on the correct direction at that exact moment.

Jane: It’s like a perfect chorus where every voice hits the right note simultaneously, rather than some voices being louder or clearer than others.

Tom: That leads perfectly into the next part of our discussion, looking at how to actually improve these results using a technique called RBA.

Improvements and RBA: Tom: We’ve established that while total diffusion is the ideal state, vanilla training methods often struggle with this consistency, leading to instability or overfitting.

Jane: The paper proposes a re-weighting scheme called Residual-based Attention, or RBA. In simple terms, it's a dynamic way of deciding which training examples are most important at any given moment.

Lu: It’s an elegant solution because it shifts the focus from treating all samples equally to giving more weight to those that are currently struggling to match the theoretical model constraints.

Meng: From an implementation perspective, RBA is designed to actively enforce that homogeneity—it pushes the system toward that ideal state where all sample gradients are aligned, which is exactly what we want for a robust AI.

Lalam: The fact that this re-weighting can accelerate the process suggests a future where AI doesn' not just learn by brute force of sheer volume of data, but by intelligently prioritizing its most challenging learning opportunities.

Tom: And the results show it is incredibly effective, reducing the time to hit ten percent L2 error by a factor of ten compared to vanilla models.

Jane: It’s like giving the model a highly efficient focus mechanism, cutting through all the noise and speeding up that final push toward total diffusion.

Tom: This brings us to our conclusion, where we synthesize all these findings before saying goodbye.

Conclusion and Synthesis: Tom: So, we've seen how "Learning in PINNs: Phase transition, total diffusion, and generalization" shows that the learning dynamics of AI are much more structured than simple steady progress.

Jane: We’ve moved from seeing those distinct phases—fitting, diffusion, total diffusion—to understanding the core mechanism that achieving a level of perfect alignment between sample gradients.

Lu: The idea of "information compression" or "binarization" is perhaps the most profound theoretical link here; it shows how AI distills complex inputs into a highly compressed, yet accurate, internal representation.

Meng: The practical implication is clear: by using RBA to accelerate entry into that total diffusion phase, we are creating significantly more stable and faster training pipelines for complex physics simulations.

Lalam: I think the ultimate impact is that we' are beginning to understand the fundamental limits and capabilities of AI when interacting with real-world physical laws, allowing us to build models that genuinely respect nature.

Tom: It’s a complete picture, moving from the abstract theory of phase transitions all the way to a practical solution in RBA.

Jane: And it' all hinges on getting that perfect "total diffusion" state where every sample is contributing equally.

Final Thoughts: Tom: Before we wrap up, I want to give each of our guests a final thought on this paper, "Learning in PINNs: Phase transition, total diffusion, and generalization."

Lu: I think the fact that these concepts map onto general theories like self-organized criticality suggests a universal principle of optimization.

Meng: For me, the clear path to faster training via RBA shows that we are ready to scale these types of scientific AI solutions commercially.

Lalam: The realization that AI can achieve a "homogenous" state of learning is truly inspiring for understanding how our machines learn.

Tom: Thank you all for sharing your insights on this fascinating research.

Jane: We hope the listeners find this discussion helpful as we wrap up the segment, and we'll be back next time with new insights into AI advancements.

Sokratis J. Anagnostopoulosa, Juan Diego Toscanob, Nikolaos Stergiopulosa, George Em Karniadakis

EPFL · Brown University, Division of Applied Mathematics · Brown University, School of Engineering, Brown University

cs.LG

Submitted: 2026-08-22

Updated: 2026-08-25

Code: https://github.com/soanagno/total-diffusion

Importance score: 80/100

The gist: * This study investigates the learning dynamics of fully-connected neural networks using Physics-Informed Neural Networks (PINNs) through the lens of gradient signal-to-noise ratio (SNR).

Key concepts

Learning Phases
The paper details three stages of AI learning: fitting, diffusion, and total diffusion. These phases describe how the network transitions through different 'states' while solving differential equations, suggesting that stability is not linear but involves distinct structural changes.
Total Diffusion
This represents a critical equilibrium state in AI learning where the signal and noise align perfectly for practical purposes. Achieving this state means training steps are consistent across all batches and suggests maximum stability for understanding physical reality.
Residual-based Attention (RBA)
RBA is a proposed re-weighting scheme that dynamically determines which training examples are most important at any moment. It shifts the focus from treating all samples equally to prioritizing those that are struggling to match theoretical constraints, thereby accelerating stable learning.

Terminology

Summary

This study investigates the learning dynamics of fully-connected neural networks using Physics-Informed Neural Networks (PINNs) through the lens of gradient signal-to-noise ratio (SNR). The research aims to provide a better understanding of the deep learning process, focusing on PINNs applications.

Theoretical Framework and Contributions

The authors analyze the training dynamics of full-batch Gradient Descent (GD) using the gradient SNR. They demonstrate that the equilibrium of the iteration-wise SNR used in Adam does not guarantee uniform learning across the sample space, which can lead to instabilities, overfitting and bad generalization. This weakness is particularly detrimental to PINNs, where the samples (collocation points) should ideally converge simultaneously to achieve a low relative error with the exact solution.

The study provides novel evidence for the existence of phase transitions proposed by the Information Bottleneck (IB) theory (fitting/diffusion phases) for PINNs trained with Adam, and reveals a third phase termed total diffusion. This phase is characterized by gradient homogeneity, where optimal convergence occurs (Figure 1).

Key contributions include:

  1. Identifying Total Diffusion: The total diffusion phase is marked by an abrupt SNR increase, uniform residuals across the sample space and the most rapid training convergence.

  2. Proposing RBA: A residual-based re-weighting scheme, termed residual-based attention (RBA), is proposed to accelerate this diffusion in quadratic loss functions. This mechanism is designed to enforce homogeneous residuals and has been shown to lead to better generalization.

  3. Linking Dynamics: The research demonstrates how the SNR and the residual diffusion coincide with information compression, which is interpreted as the binarization of activations.

Analysis of Training Dynamics

The optimization process involves minimizing a quadratic loss function:

L(theta) = 1 over n sum i in X R(f(x i; theta), y i) squared

Using Stochastic Gradient Descent (SGD), the update rule is given by theta t+1 = theta t - eta times (grad theta L + epsilon B). The core metrics used are:

  • Signal-to-Noise Ratio (SNR): Defined as SNR = E[grad theta LB] squared / Var(grad theta LB) squared. The numerator represents the gradient signal, while the denominator is the gradient noise or variation.

  • Gradient Homogeneity: A set of samples X satisfies gradient homogeneity when the gradients of all its subsets H i with fixed size m are approximately equal to the true gradient. This condition implies an ideal gradient uniformity among samples.

The analysis shows that:

  • When SNR > 1, the batch gradients become highly homogeneous.

  • When SNR 1, the optimizer is unable to make meaningful progress due to noise dominating the signal.

Residual Homogeneity and RBA

The study notes that while a stationary solution requires that, for each parameter j, sum i=1 n R i d R i over d theta j = 0, this condition is insufficient to guarantee that the residual magnitudes R are homogeneous. To achieve the required low relative L2 error, residual homogeneity must also hold.

The RBA method addresses this by dynamically adjusting the contribution of individual samples based on their residual history. The RBA calculates weights lambda i using a moving average of the relative residuals:

lambda i R i to L = (lambda i times R i(theta, x)) squared

This mechanism drives the system toward an equilibrium where lambda i about 1, and the residuals will be approximately homogeneous across the sample space X (Eq. 20).

Information Compression and Binarization

The paper investigates information compression by analyzing I(X; T) is given by:

I(X; T) = H(T) - H(TX)

The study finds that although there is some lossy compression after total diffusion, most of the information is available across all layers from the early iterations.

The findings regarding activation saturation include:

  • During the abrupt transition into total diffusion, the weights rapidly increase, causing saturation of the neuron activations.

  • The middle layers are observed to be where the middle layers appear to saturate more than the last layers, which can be interpreted by an encoding-decoding process.

Experimental Results (PINN Benchmarks)

The theoretical framework was tested on four standard PINN benchmarks: Allen-Cahn, Helmholtz, Burgers, and lid-driver cavity flow.

  • Convergence: It is during the total diffusion that the models experience the steepest convergence of the relative L2 error.

  • RBA Performance: The RBA models consistently outperform the vanilla model. For example, in Allen-Cahn, for a 10% L2 error, vanilla takes 100k iterations while RBA takes 10k, thus RBA is 10 times faster than vanilla.

  • Stability and Homogeneity: The results confirm that the total diffusion phase is linked to gradient homogeneity. In the lid-driven cavity flow (Fig. 9), the RBA model generalizes with a final L2 error of about 6%, while the vanilla model fails due to heterogeneous residuals on the right boundary, which hinder the information flow.

The study concludes that recognizing phase transitions could refine ML optimization strategies for improved generalization.

Improvements for AI systems

Based on a meticulous analysis of the research presented in Learning in PINNs: Phase transition, total diffusion, and generalization, I have developed several highly specific improvements that can be integrated into AI optimization frameworks, particularly those utilizing Physics-Informed Neural Networks (PINNs) or similar non-convex objectives.

The following improvements are designed to mitigate the limitations of standard optimizers (like vanilla Adam) in achieving optimal learning trajectories by enforcing the principles of Gradient Homogeneity and accelerating the transition into Total Diffusion.


The Improvement: Modify the standard quadratic loss function L(theta) = sum i=1 N R(f(x i; theta)) squared by dynamically re-weighting individual sample contributions based on their residual history. This is achieved by introducing a weight lambda i for each sample x i, where lambda t is calculated as a moving average of the relative residuals:

lambda t+1 = gamma lambda t + eta*(R(x i) / R infinity) squared over d R / d x i

where gamma is a decay factor, eta* is the learning rate, and R infinity is the the maximum residual magnitude.

The Capability: This mechanism ensures that samples with consistently larger residuals (those failing to satisfy physical constraints) have a disproportionately higher influence on subsequent optimizer steps. By forcing these influential weights toward lambda i about 1, this system actively drives the network towards gradient homogeneity across the entire sample space X. This results in:

  • Uniform Satisfaction: The model avoids sub-optimal solutions where certain regions of the domain are underfitted while others are overfitted.

  • Accelerated Convergence: Demonstrates a significant reduction in training time (e.g., 10x faster convergence compared to vanilla models).

The Improvement: Implement a real-time monitoring system to track the batch-wise Signal-to-Noise Ratio (SNR = E[grad theta L] squared / std[grad theta L] 2). The system is programmed to detect the onset of Total Diffusion—the abrupt increase in SNR coupled with decreasing gradient noise.

Upon detection, the system triggers a specialized Total Diffusion Acceleration protocol:

  • The learning rate eta is temporarily adjusted to maximize the deterministic component of the step, ensuring that updates are highly consistent across batches.

  • The optimizer prioritizes minimizing variance in batch gradients, effectively stabilizing the stochastic direction.

The Improvement: Integrate a mechanism to monitor the saturation of neuron activations using a Relative-Binning strategy across all layers. The system tracks both the percentage of saturated (binary) activations and the parameter norm theta.

When significant saturation is detected—especially in middle layers—the system interprets this state as achieving an efficient, compressed representation of input information. This feedback loop allows for targeted structural adjustments or loss function re-weighting to reinforce this binarization.

The improved AI system is not merely a training algorithm; it is an Adaptive, Self-Optimizing Learning Framework. It can:

  1. Dynamically Adapt: Adjust its internal weighting (lambda i) to ensure the physical constraints are met uniformly across the entire domain, rather than just averaging them.

  2. Predict and Accelerate: Identify when the learning dynamics transition into a highly efficient (Total Diffusion) state, allowing for immediate acceleration of training.

  3. Diagnose Structure: Provide detailed feedback on how information is being compressed and distributed across layers, enabling deeper architectural tuning to maximize generalization and stability in non-convex problems.

Sources

Related papers