Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs

arXiv:2607.24726 · cs.LG, cs.NA, math.NA · Submitted 2026-07-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs".

Jane: The paper was written by Justin Sirignano, Konstantinos Spiliopoulos and Samuel Cohen from University of Oxford and Boston University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the scientific machine learning world, and it's called "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs." Jane, I have to say, the title alone gets me excited because this is addressing a question that's been nagging at the field for years.

Jane: Oh absolutely, Tom. And for our listeners who might not be deep in the math weeds, let me break down what these acronyms mean. DGM stands for Deep Galerkin Method, and PINN stands for Physics Informed Neural Networks. These are both ways of using neural networks to solve partial differential equations, which are basically the mathematical rules that govern everything from fluid flow to heat distribution to how options are priced in finance.

Tom: Right, and the key word in that title is "nonlinear." Because if you're solving a simple linear equation, there's been some recent work showing these methods work. But nonlinear equations are where the real world lives, and that's where things get messy.

Jane: Exactly. And the reason it's been so hard to prove these methods work for nonlinear problems is that the objective function—the thing the neural network is trying to minimize—is what mathematicians call non-convex. That means there are lots of valleys and peaks, and a training algorithm might get stuck in a local valley that isn't actually the solution you're looking for.

Tom: So for years, people have been using these methods in practice, getting great results, but there was this nagging theoretical question: could the neural network just be converging to a wrong answer and we wouldn't know it?

Jane: Precisely. And this paper from Sirignano, Spiliopoulos, and Cohen at Oxford and Boston University finally closes that gap for a whole class of nonlinear problems. They prove that if you train long enough and make the network wide enough, the neural network will actually converge to the true solution of the partial differential equation.

Tom: And that's not just a theoretical nicety. This is the kind of result that makes people trust the tools they're already using in industry. Lu, you're the researcher here—how big a deal is this?

Lu: It's genuinely significant. The authors had to develop an entirely new mathematical approach because the previous proof techniques, which worked for linear equations, completely break down when you introduce nonlinearity. The kernel operator that governs the training dynamics becomes time-dependent, which is a whole different beast to analyze.

Jane: And they proved it in a really elegant way. They showed that as the number of hidden units in the network grows, the training process converges to a limit described by a nonlinear, non-local partial differential equation. Then they proved that this limit equation drives the residual—the error—all the way down to zero.

Tom: So the math checks out. But Meng, from your perspective as someone who actually builds these systems, what does this mean for practitioners?

Meng: Honestly, it means we can stop crossing our fingers. When I train a PINN to model some physical system, I'm always wondering if the answer I get is actually right or if it's just a local minimum that looks plausible. This paper gives us theoretical backing that with enough width and enough training time, we're converging to the real solution.

Tom: And that's the kind of confidence that lets you deploy these methods in safety-critical applications. We're going to dig into the actual proof strategy next, but first, let's just appreciate that this paper answers a question that's been open for years.

Jane: And the implications go beyond just PDEs. This is about understanding when neural network optimization actually works, which is fundamental to so much of modern machine learning.

Tom: Stay with us, because in the next segment we're going to break down how they actually proved this result, and trust me, it involves some clever mathematics that I think will surprise you.

Summary: Tom: Welcome back. We're continuing our discussion of "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs," and now we want to get into the meat of how they actually proved this thing.

Jane: Right, so the core challenge is that when you train a neural network to solve a PDE, you're minimizing the PDE residual—basically how far the network's output is from actually satisfying the equation. And this residual function is non-convex in the network parameters, so in principle, gradient descent could get stuck.

Tom: And the authors' approach to this is what I find really clever. They use something called clipped gradient descent, which is exactly what it sounds like—you clip the gradient to keep it from exploding during training. And they show that as the network gets wider, the effect of this clipping disappears, but it's crucial for getting the mathematical bounds they need.

Lu: That's right. The clipping is a practical tool that also serves a theoretical purpose. It lets them establish uniform bounds on the network's behavior during training, which is essential for the convergence proof. And importantly, the clipping doesn't change the limit dynamics—as the network width goes to infinity, the clipping vanishes.

Jane: Then they prove that the finite network converges to a limit described by a nonlinear, non-local PDE as the number of hidden units grows. This is the mean-field limit, which is a standard technique in statistical physics and has been applied to neural networks recently.

Tom: And this limit PDE is where the real action happens. It describes how the neural network output evolves during training, and it involves this kernel function that depends on the network itself. That's what makes it nonlinear and hard to analyze.

Meng: So the kernel is like the neural tangent kernel, right? The thing that describes how changes in one part of the network affect the output at another point?

Lu: Exactly. But in the linear case, that kernel is fixed—it doesn't change during training. Here, because the PDE is nonlinear, the kernel depends on the current state of the network, so it's evolving over time. That's the technical challenge that required a new approach.

Jane: And the way they handle it is by first proving that the objective function—the PDE residual—is monotonically decreasing during training. That's a key step because it gives them control over the system.

Tom: Monotonically decreasing means the error never goes up, only down. Which sounds obvious for gradient descent, but with non-convex objectives, it's not guaranteed. The gradient descent could overshoot or oscillate.

Jane: Right, but they show that with the right setup, the residual does monotonically decrease. Then they use a result called Barbalat's Lemma to show that the derivative of the objective function must go to zero, which means the system reaches a fixed point.

Lu: And then the crucial step: they characterize what those fixed points look like. They prove that any fixed point must have zero PDE residual, which means it's actually a solution to the original equation.

Meng: So the fixed points are the good points. There's no trap where the network gets stuck at a wrong answer.

Lu: Precisely. And that's the heart of the global convergence result. They show that the training dynamics must converge to a fixed point, and all fixed points are solutions. So the network converges to a solution.

Tom: And they also prove uniqueness, so even if you start from different random initializations, you end up at the same place—the actual solution to the PDE.

Jane: Which is remarkable when you think about it. The objective function is non-convex, there could be many local minima in principle, but the structure of the problem forces the dynamics to the global minimum.

Meng: And this isn't just for some toy equation. They handle a class of semi-linear PDEs, which covers a lot of real-world applications.

Tom: We're going to get into what this means for practical applications and the improvements this enables in the next segment. But I think it's worth pausing to appreciate the mathematical achievement here.

Jane: Absolutely. This is one of those results that makes you look at neural network training differently. It's not just a black box that happens to work—there's real structure that guarantees convergence.

Improvements: Tom: We're back, still talking about "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs." And now we want to talk about what this paper actually improves and enables in practice.

Jane: Right, because a theoretical result like this isn't just about satisfying mathematicians. It has real implications for how we use these methods in science and engineering.

Meng: For me, the biggest practical improvement is confidence. When I'm training a PINN to model, say, fluid dynamics or heat transfer, I now have theoretical backing that the method converges to the true solution. That means I can use these tools in situations where getting the answer wrong has real consequences.

Tom: Like in medical applications or aerospace engineering, where you can't afford to have the neural network converge to a plausible-looking but wrong answer.

Lu: And there's another improvement I find really exciting. The paper establishes uniform-in-time error bounds for the finite network. That means for a network with a fixed number of hidden units, you can bound how far it is from the true solution, and that bound doesn't degrade as training time goes to infinity.

Jane: That's a big deal. In many mean-field analyses, you can prove convergence in the limit, but you can't say anything about how fast or how uniformly. Here, they show that for any desired accuracy, you can choose a network width and training time that achieves it.

Meng: So it's not just "eventually it works." It's "here's how big your network needs to be and how long you need to train to get within this error tolerance."

Lu: Exactly. And they also show that the limits commute—you can either take the network width to infinity first and then training time, or the other way around, and you get the same answer. That's not always true in these kinds of analyses, and it's a sign that the result is robust.

Tom: And there's something else I want to highlight. The proof technique itself—the way they handle the time-dependent kernel—could be useful beyond just this specific problem.

Lu: That's a great point. The mathematical tools they developed for analyzing nonlinear, non-local PDEs that arise from neural network training could apply to other optimization problems in machine learning. It's a contribution to the broader theory of why neural networks work.

Jane: And let's not forget the practical side of the improvements. The paper shows that gradient clipping, which is already a standard technique in deep learning, doesn't hurt the convergence guarantees. That's reassuring because clipping is something practitioners already do for stability.

Meng: Right, and the fact that they use it in the proof means the theory matches what people actually do in practice. There's no gap between the idealized algorithm and the real one.

Tom: So what does this mean for the future? Where does this leave the field?

Lu: I think it opens the door to extending these results to even more general classes of PDEs. This paper handles semi-linear equations, but there are fully nonlinear equations and systems of equations that are still open.

Jane: And there's also the question of stochastic gradient descent, where you sample points randomly during training. This paper focuses on full gradient descent, but the stochastic version is what's actually used in high-dimensional problems.

Meng: So there's still work to do, but this paper provides the foundation. It's the kind of result that makes you trust the methods enough to push them further.

Tom: And that's what we're going to wrap up with in our final segment—the big picture of what this paper means for the world.

Conclusion: Tom: Alright, we're in the final stretch of our discussion on "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs." Let's pull it all together.

Jane: So to recap what we've learned: this paper proves that for a class of semi-linear PDEs, neural networks trained with gradient descent to minimize the PDE residual will converge to the true solution. No more worrying about getting stuck in local minima.

Tom: And the proof is genuinely elegant. They show the training dynamics converge to a limit described by a nonlinear PDE, they prove that PDE drives the error to zero, and they characterize all fixed points as actual solutions.

Lu: The key insight is that even though the objective function is non-convex, the structure of the PDE problem forces the dynamics to the global minimum. The time-dependent kernel, the monotonicity of the residual, the characterization of fixed points—it all fits together.

Meng: And from a practical standpoint, this means we can use DGM and PINN methods with confidence in applications where accuracy is critical. The paper also gives us uniform error bounds, so we know how big a network we need for a given accuracy.

Jane: And it's not just about this specific class of PDEs. The mathematical techniques developed here could be applied to other problems in machine learning theory.

Tom: I think the biggest takeaway is that these methods aren't just heuristics that happen to work. There's real mathematical structure guaranteeing their success. That's the kind of result that changes how people think about a field.

Lu: And it opens up new questions. Can we extend this to fully nonlinear PDEs? To systems of equations? To stochastic gradient descent? Each of those is a new challenge, but this paper provides the blueprint for how to approach them.

Meng: I'm also thinking about the practical implications for high-dimensional problems. The DGM method was designed to handle the curse of dimensionality, and now we have theoretical backing that it converges. That could make it more attractive for financial applications, where you're dealing with many variables.

Jane: And for physics and engineering, where you're modeling complex systems with many interacting components. Knowing that the neural network will find the right solution is huge.

Tom: Well, I think we've given this paper the attention it deserves. "Global Convergence of DGM and PINN Algorithms for Solving Nonlinear PDEs" by Sirignano, Spiliopoulos, and Cohen—a result that closes a long-standing gap in the theory of scientific machine learning.

Jane: And it's a beautiful example of how rigorous mathematics can provide the foundation for practical tools. We're going to be seeing more work building on this, I'm sure.

Tom: Thanks to Lu and Meng for joining us today, and to all our listeners for tuning in. We'll be back with another paper soon, so until then, keep solving those equations.

Jane: Take care, everyone.

University of Oxford · Boston University

cs.LG, cs.NA, math.NA

Submitted: 2026-07-27

Updated: 2026-09-24

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

Key concepts

Deep Galerkin Method (DGM) and PINN
These are methods that use neural networks to solve partial differential equations (PDEs). PDEs are mathematical rules governing physical phenomena, such as fluid flow or heat distribution. DGM and PINN are two ways of applying neural networks to find solutions to these equations.
Nonlinear PDEs
These are complex partial differential equations where the relationship between variables is not simple or straight. Solving these is difficult because they model messy, real-world systems, unlike simpler linear equations.
Global Convergence
In this context, it means proving that the training process of a neural network will always find the true solution (the global minimum), rather than getting stuck at an incorrect but plausible local minimum.
PDE Residual
This is the function minimized during training. It measures how far the neural network's output is from actually satisfying the original partial differential equation. The goal of training is to drive this residual down to zero.

Terminology

Summary

Summary

This paper proves the global convergence of the Deep Galerkin Method (DGM) and Physics Informed Neural Networks (PINNs) when trained with gradient descent to solve a class of semi-linear partial differential equations (PDEs). The authors address a longstanding question regarding the mathematical foundations of these algorithms: because the PDE residual objective function is non-convex, a trained neural network might, in principle, converge to a local minimizer that is not the PDE solution. The paper establishes that, for a class of semi-linear PDEs that are nonlinear in the solution and its first derivative, neural networks trained with gradient descent to minimize the PDE residual will converge to the PDE solution.

The PDE considered is of the form:

[

Au u = g, x in, u = 0, x in d,

]

where R n, Au = L + f(u, u x), f(v, w): R times R n to R, and L is a uniformly linear elliptic operator:

[

Lu = -sum i,j=1 n d over d x i [a ij(x) d u over d x j] + sum i=1 n c i(x) d u d x i - d(x)u.

]

The neural network approximation is:

[

Q N(x; theta) = eta(x)N-beta sum i=1 N c i sigma(w i x + b i),

]

where eta(x) is a smooth function vanishing on the boundary d with positive normal derivatives, N-beta is a normalization factor with 1 over 2 < beta < 1, sigma(times) is a nonlinear activation function, and the trainable parameters are theta = (c i, w i, b i) i=1 N.

The objective function to be minimized is the PDE residual:

[

J N(theta) = 1 over 2 integral (A Q N Q N(x; theta) - g(x)) squared mu(dx),

]

where mu(dx) is a probability distribution (assumed to be uniform on in Assumption 5). The training uses clipped gradient descent with learning rate alpha N = alpha N 2 beta-1.

The main results are as follows:

  1. Large N limit (Theorem 6.4): The neural network Q N(y; theta(t)) trained with clipped gradient descent converges almost surely as N to infinity to the solution of a nonlocal, nonlinear PDE:

[

d Q(t,x) over d t = -alpha integral U Q(t)(x,y) (A Q(t) Q(t,y) - g(y)) mu(dy),

]

with initial condition Q(0,x) = 0, where the kernel is:

[

U Q(t)(y,x) = grad theta eta(y)c sigma(wy+b) times D Q(t)[grad theta eta(x)c sigma(wx+b)], mu 0.

]

The convergence holds for i at most 2: x in d x i Q N(t,x) - d x i Q(t,x) to 0 almost surely on t in [0,T] for arbitrary T, and consequently t in[0,T] Q N(t) - Q(t) H squared to 0 almost surely.

  1. Convergence as t to infinity (Theorem 5.11): The limit neural network Q(t,x) converges to the solution u(x) of the PDE (1) in H 1 as t to infinity:

[

t to infinity Q(t) - u H 1 = 0.

]

  1. Convergence of the objective function (Theorem 5.10): The PDE residual R(t,x) = A Q(t) Q(t,x) - g(x) converges to zero in L squared as t to infinity, meaning the objective function J(t) to 0 and the neural network converges to a global minimizer:

[

t to infinity J(t) = t to infinity R(t) L squared squared = 0.

]

  1. Uniform-in-time error bounds for finite N (Theorem 7.1): For any epsilon > 0, there exists a t at least 0 and an N 0 < infinity such that:

[

E[J N(theta(s))] at most epsilon,

]

for all s at least t and N at least N 0. Consequently, N to infinity t to infinity E[J t N] = 0, which is a stronger result than typical mean-field analysis (which only establishes t to infinity N to infinity E[J N(theta(t))] = 0).

The proof strategy involves several key steps:

  • Fixed point characterization (Section 4): The authors prove that any fixed point of the limit PDE must be a global minimizer with zero PDE residual. Specifically, Theorem 4.4 shows that if (R,Q) in L squared times H 0 1 satisfies the fixed point equation, then R L squared = 0. This uses the fact that the linear span of eta(x) sigma(w times x + b) w,b is dense in H(0) squared = H squared H 0 1, and Assumption 3 (surjectivity of the operator D Q: H 0 1 to L squared, which is satisfied if f v(v,w) at least 0).

  • Monotonicity and a priori bounds (Section 5): Lemma 5.1 establishes that J(t) is monotonically decreasing in t, with t at least 0 R(t, times) L squared at most C < infinity. Lemma 5.4 proves uniform boundedness of Q(t) in H squared: Q(t) H squared at most C 2 < infinity. Lemma 5.7 establishes uniform bounds on d Q over d t, d Q x i over d t, and d Q x i x j over d t.

  • Barbalat's Lemma application (Lemma 5.8): The authors prove that t to infinity R(t), S Q(t) R(t) = 0 by showing that dJ over dt = -alpha R(t), S Q(t) R(t) is uniformly continuous in time.

  • Subsequence convergence (Lemma 5.9): For any sequence t n to infinity, there is a subsequence where R(t n k) weakly converges in L squared and Q(t n k) strongly converges in H 1, with the limit satisfying R, S Q R = 0.

  • Large N convergence (Section 6): The proof of Theorem 6.4 requires establishing a priori bounds on the parameter evolution (equation (57): theta i(t) - theta i(0) at most CN beta-1+ gamma with gamma + beta < 1), uniform bounds on the pre-limit PDE residual (Lemma 6.1), and showing that the clipping function's effect vanishes as N to infinity (Lemma 6.2). Lemma 6.3 bounds the difference between the pre-limit kernel U t N and the limit kernel U Q(t).

The key technical challenges addressed include: the non-convexity of the objective function (which does not convexify in the overparametrized limit for nonlinear PDEs), the time-dependence of the kernel operator S Q(t) (whose eigenvalues and eigenfunctions change over time, unlike the linear case in [2]), the lack of a spectral gap (positive eigenvalues accumulate at zero), and the need for uniform-in-time bounds on Q(t,x) and the PDE residual.

The paper concludes with a remark that the clipping of the gradient, while important for establishing the necessary a priori bounds for the analysis, has an effect that fades away as N to infinity: the clipping does not appear in the limit training dynamics.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

  • Improvement: Implement the clipped gradient descent training algorithm (with clipping parameter γ satisfying γ + β < 1) as described in Section 3, rather than standard gradient descent.

  • What the improved system can do: Guarantee that the trained neural network converges to the true PDE solution (not just a local minimum) for semi-linear elliptic PDEs of the form Au = g with nonlinearities in the solution and its first derivative. This eliminates the risk of the network converging to a spurious local minimizer that does not solve the PDE.

  • Improvement: Use the specific architecture Q N(x; θ) = η(x)N(-β) Σ c i σ(w i·x + b i) where η(x) vanishes on the boundary ∂Ω but has positive normal derivatives there, with normalization factor β ∈ (1/2, 1).

  • What the improved system can do: Automatically satisfy Dirichlet boundary conditions exactly, eliminating the need for boundary loss terms. This improves training stability and ensures the solution is in H01 (the correct function space), which is essential for the convergence proofs.

  • Improvement: Set the learning rate as α N = α·N(2β-1) (where α is a fixed positive constant) and apply element-wise smooth clipping to the gradient (with clipping function φ N that is identity on [-N γ, N γ] for γ + β < 1).

  • What the improved system can do: Maintain uniform-in-time bounds on the PDE residual and neural network parameters, preventing training instability or divergence even as the network width N grows. The clipping effect vanishes as N→∞, so it doesn't compromise the final solution accuracy.

  • Improvement: Use the PDE residual objective J N(θ) = (1/2)∫(A Q N Q N(x;θ) - g(x))2 μ(dx) with μ being a uniform probability measure on Ω.

  • What the improved system can do: Achieve mesh-free training that scales to high-dimensional PDEs (addressing the curse of dimensionality). The uniform measure ensures the residual is minimized uniformly across the domain, not just at collocation points.

  • The system can solve semi-linear elliptic PDEs with nonlinear terms f(u, ∇u) where f and its first two derivatives are bounded, and the linear operator L is uniformly elliptic with positive eigenvalues.

  • It guarantees that as the number of hidden units N→∞ and training time t→∞, the neural network converges to the unique weak solution in H01, regardless of the non-convexity of the loss landscape.

  • The system achieves lim N→∞ lim t→∞ J N(θ(t)) = lim t→∞ lim N→∞ J N(θ(t)) = 0, meaning the order of taking the network-width limit and training-time limit does not matter. This is a strong guarantee that the finite-width network actually converges to the PDE solution as training progresses.

  • For any ε > 0, the system can guarantee that there exists a finite training time t and network width N0 such that the expected PDE residual is less than ε for all subsequent training times and all larger network widths. This means the system won't forget the solution or oscillate during extended training.

  • The system can handle the case where the Neural Tangent Kernel (NTK) depends on the current solution Q(t) and its gradient (i.e., the kernel is not fixed but evolves during training). This is crucial for nonlinear PDEs where the linearization operator D Q changes as the solution evolves.

  • The system provides rigorous mathematical backing for why Physics-Informed Neural Networks and Deep Galerkin Methods work for nonlinear PDEs, addressing the longstanding open question about whether these methods can get stuck in local minima. This makes the system suitable for safety-critical applications (e.g., financial modeling, fluid dynamics, medical imaging) where unreliable numerical solutions are unacceptable.

  • By monitoring the monotonic decrease of the objective function J(t) (which is proven to be monotonically decreasing), the system can detect if training is not progressing as expected and flag potential issues, since any increase in the residual would indicate a violation of the theoretical guarantees.

  • The system can handle PDEs of the form -∇·(a(x)∇u) + c(x)·∇u - d(x)u + f(u, ∇u) = g(x) where f can depend nonlinearly on both the solution and its gradient, provided f v ≥ 0 (which ensures the linearized operator is surjective). This covers a wide range of practical applications including nonlinear heat transfer, reaction-diffusion systems, and Hamilton-Jacobi-type equations.

Related papers