From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs
summary
The gist
regularized quasi-Newton methods, which provide mechanisms for stabilizing secant models, and self-concordant methods, which provide local-metric rules for curvature-dependent step selection.
In short
The episode discusses 'SCORE,' a new quasi-Newton optimizer for Physics-Informed Neural Networks (PINNs). The authors adapted self-concordance theory, traditionally used for convex problems, to handle the complex, non-convex optimization landscape of PINNs. SCORE achieves significantly higher accuracy across multiple challenging PDEs by improving the optimization process without altering the underlying physics objective.
Key concepts
- PINNs
- Physics-Informed Neural Networks are a type of network designed to solve partial differential equations (PDEs). They achieve this by incorporating the physical laws directly into the network's loss function, which guides the learning process.
- Self-Concordance
- This is a mathematical property describing how controlled and predictable a function's curvature changes. It is useful in optimization because it allows for more reliable step size calculations, helping practitioners trust the magnitude of their optimization steps.
- Quasi-Newton Method
- These methods approximate the curvature (Hessian matrix) of a function without having to compute the full, complex matrix at every step. This makes them computationally scalable and efficient for large-scale optimization problems.
- SCORE
- SCORE is the name of the new optimizer discussed. It combines self-concordance principles with a quasi-Newton method to stabilize curvature estimates in non-convex PINN loss functions, leading to higher accuracy.
Terminology used across episodes
This episode discusses
- From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs · Paper Radio
- Lightweight Geometric Adaptation for Training Physics-Informed Neural Networks
- TINNs: Time-Induced Neural Networks for Solving Time-Dependent PDEs
- Non-Convex Self-Concordant Functions: Practical Algorithms and Complexity Analysis
- Curvature-Aware Optimization for High-Accuracy Physics-Informed Neural Networks
- Dual Natural Gradient Descent for Scalable Training of Physics-Informed Neural Networks
- Adam: A Method for Stochastic Optimization
- Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm · Paper Radio
The paper
From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs · Read on arXiv
Chenhao Si, Kang An, Shiqian Ma, Ming Yan
The Chinese University of Hong Kong, Shenzhen · Rice University · Johns Hopkins University
Physics-informed neural networks (PINNs) often require high-accuracy quasi-Newton refinement to obtain reliable partial differential equation solutions, but their residual objectives can exhibit indefinite, nearly singular, and poorly scaled local curvature. Regularized quasi-Newton methods provide established mechanisms for stabilizing secant models, while self-concordant methods provide local-metric rules for curvature-dependent step selection. Building on these two lines of work, we propose SCORE, a self-concordance-inspired quasi-Newton method with decrement-coupled shifted secant geometry for PINN training. Its distinguishing mechanism is that a single quasi-Newton decrement computed from the learned inverse metric jointly determines a strong-Wolfe-tested candidate step and an adaptive shift used to define the next secant geometry. The shifted displacement represents the action of an averaged shifted metric along the accepted step, while requiring neither Hessian construction nor Hessian-vector products. Under a local spectral-equivalence condition, we show that the quasi-Newton decrement and candidate step remain comparable to their counterparts in a positive shifted metric, and recover the normalized self-concordant rule in the matched-metric case. Strong Wolfe acceptance, fallback line search, and standard curvature safeguards provide globalization without modifying the underlying PINN objective. Experiments on the viscous Burgers, Kuramoto--Sivashinsky, Korteweg--de Vries, and complex Ginzburg--Landau equations show that SCORE attains lower final errors than the tested BFGS and self-scaled Broyden baselines. The Burgers ablation further indicates that shifted curvature stabilization and decrement-based step selection make complementary contributions to high-accuracy refinement.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs".
Jane: The paper was written by Chenhao Si, Kang An, Shiqian Ma and Ming Yan from The Chinese University of Hong Kong, Shenzhen and Rice University and Johns Hopkins University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Alright, welcome back to the show, everyone. Today we're digging into a paper that just hit arXiv with a title that's a mouthful: "From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs."
Jane: And honestly, Tom, that title packs a lot in. We've got self-concordance, quasi-Newton methods, and PINNs all in one headline. For our listeners who might not live and breathe optimization theory, let's break that down.
Tom: Please do, because I was staring at "self-concordant" for a solid minute. That's a term I don't hear every day.
Jane: So, self-concordance is a property of certain mathematical functions that tells you how wild their curvature can get. If a function is self-concordant, the way it bends changes in a very controlled, predictable way. That's super useful for optimization because you can trust your step sizes more.
Tom: And PINNs, of course, are physics-informed neural networks. Those are the networks that learn to solve partial differential equations by baking the physics right into the loss function.
Jane: Exactly. And the authors here are Chenhao Si, Kang An, Shiqian Ma, and Ming Yan, out of CUHK Shenzhen, Rice, and Johns Hopkins. They're tackling a real pain point: training PINNs to really high accuracy is notoriously hard.
Tom: Hard how? I mean, we've talked about PINNs before on this show, and they seem to work pretty well.
Jane: They do, but there's a catch. When you're trying to get a PINN to produce a really precise solution, the optimization landscape becomes this nasty, bumpy, ill-conditioned mess. The curvature information you need to make good progress can be indefinite, nearly singular, or just badly scaled.
Tom: So the math equivalent of trying to walk downhill in the dark with a broken flashlight.
Jane: That's exactly it. And the authors' big idea is to borrow tools from self-concordance theory, which was designed for nice convex problems, and adapt them to this messy nonconvex PINN world.
Tom: So they're taking a tool from the clean, tidy convex world and forcing it to work in the wild west of nonconvex optimization.
Jane: Right. And they call their method SCORE. It's a quasi-Newton method, which means it approximates the curvature without computing the full Hessian matrix, and it uses this self-concordance-inspired rule to decide how big a step to take and how to stabilize the curvature estimate.
Tom: And the payoff? They're reporting some seriously impressive error reductions on benchmark equations like Burgers and Kuramoto-Sivashinsky.
Jane: We're talking about final errors that are several times smaller than what standard methods like BFGS achieve. For the Burgers equation, they got a relative L2 error of two point two five times ten to the minus nine.
Tom: That's nine decimal places of accuracy. That's not just a small improvement, that's a whole different league.
Jane: And the exciting part is that the method doesn't change the underlying physics objective at all. It's purely a better optimizer, which means it could be dropped into existing PINN workflows.
Tom: So before we get into the nitty-gritty of how SCORE actually works, I want to know one thing. Is this a paper that's going to matter to people who aren't optimization theorists?
Jane: Oh, absolutely. Anyone who's ever trained a PINN and watched it plateau at a mediocre error level is going to care about this. And that's a lot of people in computational science right now.
Tom: Then let's get into the details. I want to understand the mechanism behind this thing, because the title alone suggests they're doing something clever with the geometry of the loss landscape.
Jane: And that's exactly where we're headed next.
Summary of the Paper: Tom: So we've established that SCORE is a new optimizer for PINNs that gets crazy good accuracy. Now let's talk about what the paper actually does under the hood.
Jane: Right. And to understand it, you have to remember what a PINN loss function looks like. It's a sum of squared residuals. You've got the PDE residual, the boundary condition residual, and the initial condition residual, all squared and added together.
Tom: Sum of squares. So it's like a least-squares problem.
Jane: It is, but here's the twist. The full Hessian of that objective has two parts. One part comes from the Gauss-Newton approximation, which is always positive. But there's a second part that depends on the residual values themselves, and that part can be positive or negative.
Tom: So the curvature can flip sign depending on where you are in the landscape.
Jane: Exactly. And when you're trying to do quasi-Newton refinement, which relies on building up a good approximation of the curvature from successive steps, that sign-indefinite part is poison. It makes your curvature estimate unreliable.
Tom: So the raw material that quasi-Newton methods use to learn the landscape is contaminated.
Jane: That's the problem they're solving. And their solution draws on something called weak self-concordance. In classical self-concordance, you have a function whose curvature changes in a controlled way. But that requires the function to be convex, which PINN losses are not.
Tom: So how do you get around that?
Jane: You add a positive shift to the curvature. Instead of working with the raw Hessian, you work with the Hessian plus some positive constant times the identity matrix. That guarantees you have a positive definite local metric to measure things with.
Tom: So you're artificially inflating the curvature to make it well-behaved.
Jane: Yes, but the trick is doing it smartly. The shift isn't fixed; it's adaptive. And this is where the cleverness comes in. They compute something called a quasi-Newton decrement, which is a measure of how big the proposed step is in the current learned metric.
Tom: And that decrement drives everything?
Jane: It drives two things. First, it determines the candidate step length. When the decrement is large, meaning you're far from the solution, you take a smaller step. When it's small, you take a bigger step. That's the self-concordant damping rule.
Tom: So the step size is calibrated to the local geometry.
Jane: And second, the same decrement determines the shift used to define the next curvature estimate. Big decrement means a stronger shift, which means more stabilization. Small decrement means less shift.
Tom: So one number controls both how far you step and how much you stabilize the next step.
Jane: Exactly. It's an elegant coupling. The metric you learned at iteration k controls the step you take, and the step you take generates the data for the next metric.
Tom: And they implement this without ever forming the Hessian. That's the scalable part.
Jane: Right. They use a self-scaled Broyden update, which is a quasi-Newton formula, and they just replace the raw gradient displacement with a shifted version. That shifted displacement represents the action of the positive local metric along the accepted step.
Tom: So they're baking the regularization directly into the secant equation.
Jane: Precisely. And the whole thing is wrapped in a strong Wolfe line search, so it's globally convergent in practice. The candidate step is tested, and if it fails, they fall back to a standard line search.
Tom: That's a robust design. You get the benefits of the self-concordant step when it works, and you have a safety net when it doesn't.
Jane: And the experiments show that both components matter. They did an ablation on the Burgers equation where they removed each piece separately, and both contributed to the final accuracy.
Tom: So it's not just one trick carrying the whole thing.
Jane: No, it's a combination. The shifted secant and the decrement-based step selection each help independently, and together they give you the best result.
Tom: Okay, I'm starting to see why this is exciting. But I want to know how it actually performs in practice. Let's talk about those experiments.
Jane: Good, because that's where the rubber meets the road.
Improvements and Experiments: Tom: So we've got the theory. Now let's talk about the numbers. What did SCORE actually achieve on these benchmark problems?
Jane: The paper tests on four equations: viscous Burgers, Kuramoto-Sivashinsky, Korteweg-de Vries, and complex Ginzburg-Landau. And the results are pretty striking across the board.
Tom: Give me the highlights.
Jane: On Burgers, SCORE gets a relative L2 error of two point two five times ten to the minus nine. The best baseline, SSBroyden, gets one point four times ten to the minus eight. So SCORE is about six times more accurate.
Tom: Six times better on a classic benchmark. That's not nothing.
Jane: And on the Kuramoto-Sivashinsky equation, which is a chaotic system, SCORE gets one point eight four times ten to the minus four, while BFGS gets five point five nine times ten to the minus four. That's a three-fold improvement on a genuinely hard problem.
Tom: Chaotic systems are brutal for PINNs because small errors get amplified over time.
Jane: Exactly. And they used a time-marching strategy where they train on consecutive time windows, passing the solution forward. SCORE consistently outperformed the baselines in the later windows, which are the hardest ones.
Tom: So the advantage grows as the problem gets harder.
Jane: That's the pattern. On the Korteweg-de Vries equation, which has soliton-like wave dynamics, SCORE gets three point five seven times ten to the minus seven, versus eight point five seven times ten to the minus seven for SSBroyden. Again, more than twice as accurate.
Tom: And the complex Ginzburg-Landau equation? That's a two-dimensional complex-valued problem.
Jane: Right, that's the most challenging one. They train a network that outputs both the real and imaginary parts of the field. SCORE gets relative L2 errors around eight point four times ten to the minus four on both components, while the baselines are in the two to two point seven times ten to the minus three range.
Tom: So a solid two to three times improvement even in that high-dimensional setting.
Jane: And importantly, the computational cost is comparable. The wall-clock time per refinement block is about the same for all three methods. SCORE isn't slower; it's just more accurate.
Tom: That's the best kind of improvement. Better results, same cost.
Jane: And they also ran an ablation on Burgers to isolate the two key components. The full SCORE method gets two point two five times ten to the minus nine. If you only use the shifted secant, you get five point three two times ten to the minus nine. If you only use the self-concordant step test, you get nine point nine nine times ten to the minus nine.
Tom: So each piece helps, but together they're better than the sum of the parts.
Jane: Exactly. And all of them beat the plain SSBroyden baseline at one point four times ten to the minus eight. So both innovations are pulling their weight.
Tom: I'm curious about the practical implications here. If I'm a researcher using PINNs, do I need to change my whole pipeline to use this?
Jane: No, and that's the beauty of it. The method is a drop-in replacement for the optimizer. You keep your network architecture, your loss function, your collocation point sampling. You just swap out the quasi-Newton optimizer for SCORE.
Tom: So it's a plug-and-play improvement.
Jane: Essentially. And the hyperparameters are kept fixed across all four benchmarks, so there's no problem-specific tuning needed. That's a big deal for real-world usability.
Tom: That's actually huge. So many optimization papers work great on one problem but fall apart when you change the equation.
Jane: And here they've shown robustness across four very different PDEs, from chaotic systems to dispersive waves to complex-valued fields. That suggests the method is capturing something fundamental about the optimization geometry.
Tom: So what's the big picture here? Where does this leave the field?
Jane: I think it points toward a future where PINN training is guided by the geometry of the loss landscape itself, rather than by generic step-size rules. And that could unlock the kind of high accuracy that makes PINNs viable for serious scientific computing.
Tom: That's a compelling vision. Let's bring in the rest of the team to get their take.
Conclusion: Tom: So we've covered the title, the method, and the results. Let's wrap up our discussion of "From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs" with the whole crew.
Jane: I'd love to hear what Lu thinks about the broader implications, because this feels like it could change how people approach PINN optimization.
Lu: I think the most exciting thing is that this is a principled approach, not just a heuristic. The authors took a deep theoretical concept, weak self-concordance, and found a practical way to realize it in a quasi-Newton framework. That's rare.
Meng: But I have to ask the engineer's question. How robust is this in practice? The paper shows great results on four benchmarks, but what about real-world problems with messy data and noisy gradients?
Jane: That's a fair concern. The paper uses clean simulation data and deterministic residual evaluations. Real-world applications often have measurement noise and more complex constraints.
Lu: But the fact that it works across four very different PDEs suggests the underlying principle is sound. And the method doesn't change the objective, so it should be compatible with existing robustness tricks like loss weighting and adaptive sampling.
Meng: And the computational overhead is minimal, right? They showed comparable wall-clock times to standard BFGS.
Jane: Right, the per-block time is essentially the same. The shifted secant construction is cheap, and the candidate step test is just a few extra evaluations.
Tom: So it's not a method that costs you anything to try. You just swap it in and see if it helps.
Meng: That's the kind of thing I like. Low risk, potentially high reward.
Lalam: I want to add a cultural perspective. This paper represents a shift in how we think about AI for science. Instead of treating the optimizer as a generic black box, we're now designing optimizers that understand the structure of the physics they're solving.
Tom: That's a nice way to put it. The optimizer becomes physics-aware, not just the network.
Lalam: And that has implications beyond PINNs. Any scientific machine learning task where the loss function has a known structure could benefit from this kind of geometry-aware optimization. Climate modeling, drug discovery, materials science, all of these rely on solving PDEs accurately.
Jane: And the accuracy gains here are not marginal. We're talking about factors of two to six in error reduction. For scientific applications where precision matters, that could be the difference between a useful simulation and a useless one.
Lu: I'd also note that the theoretical framework could inspire further work. Now that we have a self-concordance-inspired quasi-Newton method, can we prove convergence rates? Can we extend it to other loss structures?
Meng: And can we make it work in a distributed setting? Some of these PDE problems are huge, and you need to parallelize the training.
Tom: Lots of open questions, but that's what makes this exciting. The paper opens a door.
Jane: It does. And I think the takeaway for our listeners is simple. If you're training PINNs and struggling to get high accuracy, this is a method worth trying. It's free, it's fast, and it works.
Tom: So with that, let's say goodbye to "From Non-Convex Self-Concordant Regularization to Scalable Quasi-Newton Training of PINNs." Great paper, great results, and we're looking forward to seeing where this line of research goes.
Jane: Absolutely. Thanks for joining us, everyone. We'll be back soon with the next paper on arXiv.
Tom: Until then, keep optimizing.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization