Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm

arXiv:2602.08515 · math.NA, cs.LG, cs.NA, cs.NE, math.OC · Submitted 2026-02-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Discussion of Summary and Scope: Tom: The authors addressed both forward problems—where we know the physics and solve for the solution—and inverse problems, where we have data but need to find unknown parameters in "Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg–Marquardt algorithm."

Jane: That wide scope is crucial. It shows this isn't just a niche optimization; it' addressing the core challenges of applying AI to real-world physical systems.

Lu: The fact they tackle both forward and inverse problems suggests a universal applicability, not just that one specific task works well in the framework.

Meng: Inverse problems are notoriously hard because we have incomplete information, so getting results there is a huge practical win for an industry.

Lalam: Being able to identify unknown parameters using shallow networks feels like democratizing scientific discovery by making complex data easier to interpret.

Discussion of Improvements and Methodology: Tom: The core improvement centers on how the authors formulated the problem, turning it into a nonlinear least-squares system that can be solved using the Levenberg–Marquardt algorithm.

Jane: That's a big technical leap. Instead of just running standard gradient descent, they’ are treating it like solving a specific mathematical system of equations.

Lu: The authors derived exact analytical expressions for the neural network derivatives to compute the Jacobian matrix, which is key to understanding how the model behaves locally.

Meng: From an implementation standpoint, calculating that Jacobian analytically sounds much faster than relying solely on general automatic differentiation tools for a shallow architecture.

Lalam: The explicit visualization of how shared weights contribute to both solution and its derivatives gives us a clearer view into the internal workings of AI systems.

Discussion of Results and Experiments: Tom: They tested this approach on several benchmark problems, including Burgers, Schrödinger, Allen–Cahn, and the three-dimensional Bratu equations.

Jane: It's reassuring to see these diverse tests because if they work on simple systems like Burgers but also tackle complex nonlinear differential operators is a huge deal.

Lu: The consistent performance across different physical laws suggests that the optimization method is robust, regardless of whether it’s a time-dependent or time-independent problem.

Meng: The results show that LM significantly outperforms BFGS in convergence speed and final loss values; that's a direct efficiency boost for computational resources.

Lalam: It feels like the combination of shallow networks and this specific optimization strategy is streamlining how we approach complex mathematical challenges in society.

Conclusion and Final Thoughts: Tom: We’ve seen how "Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg–Marquardt algorithm" offers a powerful alternative to deep models.

Jane: It really proves that an intelligent optimization strategy is just as important as the architectural complexity itself for achieving high accuracy in solving PDEs.

Lu: The paper has opened up a new paradigm where computational efficiency doesn' and depth are not necessarily coupled, which is a massive theoretical win.

Meng: I think this means we can deploy complex models on less powerful hardware while maintaining high accuracy, which is highly practical for real-world deployment.

Lalam: It allows us to build more elegant, efficient solutions to share with the world, making advanced science accessible and sustainable.

Tom: That’s a great way to wrap up our discussion on "Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg–Marquardt algorithm."

Lu: It’s a genuinely elegant solution.

Meng: I'm already thinking about how this will scale.

Lalam: This is a significant step forward for AI culture.

math.NA, cs.LG, cs.NA, cs.NE, math.OC

Submitted: 2026-02-09

Updated: 2026-08-25

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 94/100

The gist: This paper investigates whether deep architectures are essential for physics-informed neural networks (PINNs) by proposing the use of shallow neural networks combined with the Levenberg–Marquardt

Key concepts

Physics-Informed Neural Networks (PINNs)
The scope of the research involves addressing two types of problems: forward problems, where physics are known to solve for a solution; and inverse problems, where data is available but unknown parameters must be identified.
Levenberg–Marquardt Algorithm
This algorithm was used because the authors reformulated the problem into a nonlinear least-squares system. It proved to be highly effective, showing significant improvements in convergence speed and final loss values compared to BFGS.
Forward and Inverse Problems
The paper tackles both scenarios. Forward problems involve solving for a solution when physical laws are known. Inverse problems involve using incomplete data to find unknown parameters within a system.

Terminology

Summary

This paper investigates whether deep architectures are essential for physics-informed neural networks (PINNs) by proposing the use of shallow neural networks combined with the Levenberg–Marquardt (LM) algorithm. By reformulating PINNs as nonlinear systems, the authors demonstrate that shallow architectures can solve both forward and inverse problems for nonlinear partial differential equations (PDEs) with high efficiency and accuracy, potentially reducing computational costs.

How it works

The researchers propose using shallow feed-forward neural networks with two hidden layers to approximate solutions. Instead of the standard approach of minimizing a composite loss function using first-order optimizers like Adam or quasi-Newton methods like LBFGS, they reformulate the training process as solving a nonlinear system of equations, or equivalently a nonlinear least-squares problem.

The Levenberg–Marquardt (LM) algorithm is employed to optimize network parameters, acting as a robust and efficient optimization method that blends the Gauss–Newton algorithm and gradient descent. To facilitate this, the authors develop an exact analytical differentiation scheme to compute the Jacobian matrix. This approach involves:

  • Rewriting network operations in a layer-wise form for clarity.

  • Deriving analytical expressions for derivatives with respect to input variables x and t.

  • Constructing the Jacobian matrix J from the partial derivatives of each residual with respect to all network parameters.

Problem Formulations

For forward problems, the objective is to approximate the solution by minimizing the PDE residual. The authors enforce initial and boundary conditions directly within the network representation using a transformation that improves training efficiency and solution accuracy.

In inverse problems, unknown parameters lambda are identified simultaneously with the solution u. The Jacobian matrix for these problems takes a specific block form to handle both network weights and PDE parameters:

  1. J 1: Derivatives of the PDE residuals with respect to the network weights.

  2. J 2: Derivatives of the data-matching residuals with respect to the network weights.

  3. J 3: Derivatives of the PDE residuals with respect to the unknown parameters lambda.

Experimental Results

The proposed method was tested on several benchmark problems to evaluate its performance. These included:

  • The Burgers equation.

  • The nonlinear Schrödinger equation.

  • The Allen–Cahn equation.

  • The three-dimensional Bratu equation.

Numerical results demonstrate that LM significantly outperforms BFGS in terms of convergence speed, accuracy, and final loss values. For the Burgers equation, LM achieved a final loss of 8.1 times 10-8, which is significantly smaller than the 1.4 times 10-3 obtained by BFGS. In the Schrödinger equation case, the shallow network attained a relative L2 error of 2.1 times 10-5, representing an improvement of nearly two orders of magnitude in accuracy while requiring more than 25 times fewer parameters compared to deeper models. For inverse problems, LM converges to the true parameter values both faster and more accurately than BFGS.

Implications of Network Depth

The study concludes that network depth and size alone do not determine the accuracy of PINNs when an appropriate optimization strategy is employed. Shallow architectures, when combined with effective second-order optimization methods, can provide accurate and computationally efficient solutions for both forward and inverse problems. This suggests that deep architectures may not be essential for high-fidelity results in a wide class of PDEs.

Improvements for AI systems

1. Implementation of Shallow-Architecture PINN Modules

  • Improvement: Replace deep, high-parameter neural architectures with parsimonious, two-hidden-layer structures (NN(a, m 1, m 2, 1)) specifically optimized for physical constraints.

  • Capability: The improved AI system can solve complex nonlinear partial differential equations (PDEs) using significantly fewer trainable parameters (up to 25x reduction) while maintaining or exceeding the accuracy of deep models. This enables high-fidelity physical simulations on edge devices and low-resource hardware with minimal memory footprints.

2. Integration of Levenberg–Marquardt (LM) Optimization Engines

  • Improvement: Shift from first-order (Adam) or quasi-Newton (LBFGS) optimizers to a second-order Levenberg–Marquardt algorithm specifically formulated for nonlinear least-squares residual minimization.

  • Capability: The AI system can achieve much higher precision, reaching loss values in the 10-7 to 10-10 range, where standard optimizers often stagnate at 10-3. This results in faster convergence and more stable training for highly nonlinear systems like the Schrödinger or Burgers equations.

3. Transition from Automatic Differentiation to Exact Analytical Derivative Schemes

  • Improvement: Incorporate the paper’s derived analytical expressions for spatial, temporal, and weight-based derivatives into the backpropagation pipeline, replacing reliance on standard automatic differentiation (AD).

  • Capability: The system can perform extremely rapid and computationally efficient Jacobian matrix evaluations. This reduces the computational overhead per iteration and provides a transparent, mathematically rigorous dependency structure between network weights and the resulting differential operators, enhancing training stability.

4. Unified Forward-Inverse Nonlinear System Solver

  • Improvement: Reformulate PINN training from simple loss minimization to solving a unified nonlinear system F(W) = 0.

  • Capability: The AI system can perform simultaneous state estimation (forward problem) and high-precision parameter identification (inverse problem). For example, it can identify unknown physical constants (e.g., diffusion coefficients or reaction rates) in real-time with extremely low absolute percentage errors (APE), making it ideal for autonomous system calibration and digital twin synchronization.

Abstract

This work investigates shallow physics-informed neural networks (PINNs) for solving forward and inverse problems governed by nonlinear partial differential equations (PDEs). By formulating PINN training as a nonlinear least-squares problem, the Levenberg-Marquardt (LM) algorithm is used to efficiently optimize the network parameters. Exact analytical expressions for neural-network derivatives with respect to the input variables are derived, revealing the relationships between the network output and its spatial and temporal derivatives and providing a clearer interpretation of the PINN architecture. These expressions are then used to derive explicit formulas for the Jacobian matrix required by LM. The proposed approach is evaluated on the Burgers, Schrödinger, Allen-Cahn, and three-dimensional Bratu equations. Numerical results show that LM substantially outperforms BFGS, L-BFGS, and Adam in convergence speed, accuracy, and final loss values. Comparisons with deeper networks further demonstrate that shallow LM-PINNs can achieve higher accuracy with substantially fewer parameters, emphasizing the importance of considering network architecture and optimization strategy jointly. The explicit analytical Jacobian also provides computational and memory advantages that are particularly relevant to large-scale PINNs. Overall, these results suggest that, for a broad class of PDEs, shallow PINNs combined with effective second-order optimization can provide accurate and computationally efficient solutions to both forward and inverse problems.

Sources

Related papers