Near-Equilibrium Propagation training in nonlinear wave systems

arXiv:2510.16084 · cs.LG, cond-mat.quant-gas, math-ph, math.MP, physics.optics, quant-ph · Submitted 2026-08-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Near-Equilibrium Propagation training in nonlinear wave systems".

Jane: The paper was written by Karol Sajnok and Michał Matuszewski from Center for Theoretical Physics, Polish Academy of Sciences and Institute of Physics, Polish Academy of Sciences.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Paper discussion segment 1: Tom: Welcome back to the show, everyone. Today we're diving into a fascinating new paper that just hit arXiv, called "Near-Equilibrium Propagation training in nonlinear wave systems." Jane, I have to say, when I first read the title I thought, okay, this is going to be dense. But the idea here is actually pretty wild.

Jane: It really is, Tom. And I think the best way to explain it is to start with the problem. You know how training a neural network normally works, right? You have to send signals backward through the network to calculate how to adjust the weights. That's backpropagation, and it's incredibly powerful but it assumes you have a perfect digital model of your system.

Tom: Right, and that's fine when you're running on a GPU. But these researchers are working with physical systems — actual wave systems, like exciton-polariton condensates. These are real physical objects made of light and matter. And when you try to train them in the real world, backpropagation falls apart because you don't have a perfect model.

Jane: Exactly. So they've taken this alternative approach called Equilibrium Propagation, which was developed a few years back. The trick is you let the physical system relax into a steady state, then you nudge it slightly toward the answer you want, and you compare the two states. That difference gives you the gradient you need to update the system. No backward pass needed.

Tom: And the new twist in this paper, "Near-Equilibrium Propagation training in nonlinear wave systems," is that they've made this work for complex-valued wave systems. That's a big deal because waves have both amplitude and phase, and they oscillate. Previous work mostly assumed you had some kind of energy function that the system was trying to minimize. But wave systems don't work that way — they're driven and dissipative.

Jane: Right, they're not relaxing to a minimum energy state. They're being pumped with energy and losing energy at the same time, and they settle into a steady oscillation. The authors call this "near-equilibrium" because it's not true equilibrium — it's a driven steady state. And they show that even in that regime, you can still extract the correct gradients.

Tom: And they don't just prove it theoretically. They actually simulate training a physical polariton system to do real tasks. We're talking XOR logic and even handwritten digit recognition on MNIST. That's pretty impressive for a physical system.

Jane: It is. And I love that they're training local parameters — like the potential landscape of the system — rather than just the connections between nodes. That's a much more realistic thing to control in a lab. You can shine a laser pattern onto the sample to create the potential you want.

Tom: So this isn't just a theoretical curiosity. It's a practical recipe for building physical neural networks that can learn in the real world. And that could be huge for energy efficiency, because you're not simulating anything — the physics does the computation for you.

Jane: Exactly. And I think the key insight is that the system itself becomes the learning machine. The gradient information is extracted from the physical response of the system, not from a digital model. That's the breakthrough here.

Tom: So what do you think, Lu? You've been working on physical neuromorphic systems for a while. Does this paper move the needle for you?

Lu: Oh, absolutely, Tom. I think the most exciting part is the generality. They've shown this works for both discrete and continuous systems, and they've derived the update rules from first principles. That means it's not just applicable to polaritons — it could work in photonic circuits, mechanical oscillators, even certain kinds of electronic systems. The framework is really quite general.

Tom: And that generality is what could make this a foundational paper. Let's keep digging into the details in the next segment.

Paper discussion segment 2: Jane: So we're back, and we're still talking about "Near-Equilibrium Propagation training in nonlinear wave systems." Tom, I want to go a bit deeper into how the nudging actually works, because I think that's the cleverest part of the whole scheme.

Tom: Yeah, let's break it down. In the first phase, you let the system evolve under the input until it reaches a steady state. You measure the output. Then in the second phase, you add a small perturbation — a nudge — that pushes the output toward what you wanted it to be. The system settles into a slightly different steady state.

Jane: And the difference between those two states, the free state and the nudged state, gives you the gradient. But here's the thing — in a wave system, you can't just add any perturbation. The authors had to be careful about how they define the nudging term. They use something called Wirtinger derivatives, which handle complex variables properly.

Tom: Right, because the wavefunction has both a real and imaginary part, and you need to treat them correctly. The nudge they add includes an imaginary factor, which corresponds to a phase shift in the driving beam. That's actually something you can do experimentally — you just shift the phase of your laser.

Jane: And that's the beauty of it. The nudging is implemented as a resonant pump in the output region, with an amplitude proportional to the error and a phase shift. That's a very clean experimental implementation. You're literally shining light on the sample to tell it what it got wrong.

Tom: Now, one thing I want to ask about — the paper mentions something called "near-equilibrium conditions." What does that mean exactly?

Jane: So in the ideal case, the system's dynamics would come from a Hamiltonian, and the equations would satisfy certain symmetry conditions. But real physical systems have dissipation and losses, which break those symmetries. The authors show that as long as the dissipation is small compared to the other energy scales, the gradient estimate is still accurate enough to train successfully.

Lu: And that's actually a really important practical point, Jane. In real experiments, you can't eliminate losses. But you can design systems where the losses are small enough that the near-equilibrium approximation holds. The paper even quantifies this — they show that training works well when the nonlinearity-to-loss ratio is around zero point three.

Tom: That's a specific number we can hold onto. And it's not just a theoretical bound — they tested it numerically. They ran the full simulation of the Gross-Pitaevskii equation, which is the equation that describes polariton condensates, and they showed that training converges reliably.

Jane: And the results are pretty striking. For the MNIST task, they got about ninety percent test accuracy. That's comparable to what you'd get from a simple fully connected network trained with backpropagation on the same data. So the physical learning is not just working — it's working as well as digital learning for this kind of task.

Meng: Can I jump in here? I'm curious about the practical side. How long does this training actually take? Because if you have to wait for the system to reach steady state for every sample, that could be slow.

Jane: Great question, Meng. The paper addresses exactly that. They estimate that for typical polariton systems, the full training procedure would take somewhere between zero point one and ten milliseconds. That's incredibly fast — orders of magnitude faster than GPU-based simulation.

Meng: Wait, really? Milliseconds for the entire training run? That's remarkable. Because on a GPU, simulating these dynamics is computationally expensive. But if the physical system does the computation itself, it's just the time it takes for the condensate to relax.

Tom: Exactly. And that's the whole point of physical neuromorphic computing. The physics is the computation. You're not simulating the wave equation — you're letting the wave equation happen.

Jane: And the authors note that this is nearly six orders of magnitude faster than numerical integration on a GPU. That's a huge practical advantage, and it's what makes this approach so compelling for real-world applications.

Lu: I'd also add that the fact they can train local potentials rather than just connection weights is a big deal for scalability. You can tile these systems into weakly coupled cells, each with its own trainable potential, and train them all in parallel. That's a path toward much larger physical neural networks.

Tom: So we've got speed, we've got scalability, and we've got a clear experimental recipe. What could go wrong? Let's talk about the limitations and what this means for the future in the next segment.

Paper discussion segment 3: Jane: Welcome back. We're still on "Near-Equilibrium Propagation training in nonlinear wave systems," and I want to talk about what this paper improves compared to what came before. Because there's been other work on Equilibrium Propagation in physical systems, but this one really pushes things forward.

Tom: Right, and I think the biggest improvement is that it works for wave systems that don't have a well-defined energy function. Previous approaches assumed you had some kind of energy landscape the system was relaxing into. But wave systems are different — they oscillate, they have phase, they're driven by external pumps.

Jane: Exactly. And the authors point out that some earlier work on oscillatory systems with EP assumed energy-gradient descent, which gives you relaxation rather than oscillation. That doesn't match how real wave systems behave. This paper handles the actual Hamiltonian dynamics, including the oscillatory behavior.

Tom: Another improvement is that they can train local parameters — like the potential at each point in space — rather than just the weights between nodes. That's a much more flexible framework because it applies to systems where you don't have well-defined nodes and connections.

Jane: And that's actually a really important point. In a continuous wave system, there aren't discrete "neurons" with "synapses." It's a continuous field. So you need a training rule that works with continuous parameters, not just discrete weights. This paper provides exactly that.

Lu: I'd also highlight the fact that they use Wirtinger derivatives, which is the right mathematical tool for complex-valued systems. Previous work on EP in complex networks assumed holomorphic functionals, which don't correspond to physical quantities like energy. This paper uses the correct framework, and that makes the derivation rigorous.

Tom: And there's a practical improvement too. The nudging is implemented as a resonant pump with a phase shift, which is something you can actually do in the lab. It's not an abstract mathematical perturbation — it's a physical beam of light.

Jane: Right. And they also show robustness to imperfections. They tested training with fixed random fluctuations in the potential, simulating structural inhomogeneities in real samples, and the training still converged. That's important because real experimental samples are never perfect.

Meng: So what about the limitations? I'm guessing the near-equilibrium condition is a real constraint. You can't have arbitrarily strong dissipation.

Jane: That's right, Meng. The method works best when dissipation is small compared to the other energy scales. If you have too much loss, the gradient estimate becomes inaccurate and training can destabilize. They showed this in their experiments — larger nonlinearity-to-loss ratios led to worse performance.

Tom: But even with that limitation, the window of operation is pretty wide. And for polariton systems, which are the ones they're targeting, that window is experimentally accessible.

Lu: I think the biggest impact of this paper is that it provides a practical blueprint for in-situ training. You don't need to simulate anything — you just need to measure the steady states and apply the update rules. That's a huge step toward building physical neural networks that can learn in the real world.

Jane: And the authors also mention that a related scheme was published after their work, which shows this is a hot area. The field is moving fast, and this paper is going to be a reference point for future work.

Tom: So let's bring in Lalam to give us the big-picture view. What does this mean for the world beyond the lab?

Conclusion: Tom: So we're wrapping up our discussion of "Near-Equilibrium Propagation training in nonlinear wave systems." Jane, can you give us a final summary of what this paper achieves?

Jane: Sure, Tom. This paper takes the idea of Equilibrium Propagation — training physical systems by comparing free and nudged steady states — and extends it to complex-valued wave systems that operate in a driven-dissipative regime. They derive the update rules rigorously, show they work in numerical simulations of polariton condensates, and demonstrate successful training on XOR and MNIST tasks.

Tom: And the key takeaway is that you can train a physical wave system to perform computation without needing a digital model. The system learns in situ, using its own physics to calculate the gradients. That's a fundamentally different approach to neuromorphic computing.

Jane: The practical implications are huge. We're talking about training times on the order of milliseconds, energy efficiency that's orders of magnitude better than digital simulation, and a clear experimental recipe using standard optical tools.

Lu: And the generality of the framework means it could apply to many different physical platforms, not just polaritons. Photonic circuits, mechanical oscillators, maybe even certain electronic systems. This is a foundational result.

Meng: From an engineering standpoint, the fact that they train local potentials rather than connection weights is really practical. It means you can use spatial light modulators and camera readouts, which are standard tools. The implementation path is clear.

Lalam: I'd add that this kind of physical learning could eventually lead to devices that learn continuously in their environment, adapting to real-time data without the energy cost of digital training. That could transform edge computing, robotics, and even scientific instrumentation where power is limited.

Tom: Well said, Lalam. And with that, we'll say goodbye to "Near-Equilibrium Propagation training in nonlinear wave systems." It's a paper that brings us one step closer to computers that learn with light and matter instead of silicon and electricity.

Jane: And we're excited to see where this line of research goes next. Thanks for listening, everyone. We'll be back with another paper soon.

Tom: Until then, keep your eyes on the arXiv. There's always something new and surprising. Take care, everyone.

Karol Sajnok, Michał Matuszewski

Center for Theoretical Physics, Polish Academy of Sciences · Institute of Physics, Polish Academy of Sciences

cs.LG, cond-mat.quant-gas, math-ph, math.MP, physics.optics, quant-ph

Submitted: 2026-08-16

Updated: 2026-08-18

Comments: 7 figures

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

Key concepts

Equilibrium Propagation (EP)
This is an alternative training method used in physical systems that does not require a perfect digital model. Instead of using backpropagation, the system is allowed to reach a steady state and then slightly nudged. The difference between these two states provides the necessary gradient required to update the system.
Near-Equilibrium State
This refers to a driven, dissipative steady oscillation where a physical system settles into a stable state by being pumped with energy while simultaneously losing it. The authors demonstrate that even in this non-minimal, steady state, the correct gradient information can be extracted for successful training.
Wirtinger Derivatives
Because wave functions possess both amplitude and phase (they are complex-valued), standard mathematical tools are insufficient. Wirtinger derivatives are the specific mathematical framework used by to correctly handle these complex variables when calculating necessary perturbations or gradients in the physical system.
Local Parameters
Unlike traditional neural networks that train connections between discrete nodes, this method allows researchers to train local parameters, such as the potential landscape of a continuous physical system. This is practical because it can be controlled by applying a laser pattern to create the desired potential.

Terminology

Summary

Summary

This paper introduces Near-Equilibrium Propagation (NEP), a generalization of the Equilibrium Propagation (EP) learning algorithm designed for training driven–dissipative complex-valued wave systems. The authors note that backpropagation, the standard training method for artificial intelligence, is difficult to implement in physical neural networks due to its reliance on an explicit computational graph and accurate adjoint dynamics, which are fragile in analog hardware. EP is presented as an alternative with comparable efficiency and strong potential for in-situ training, but previous EP implementations were limited. The paper states: "We extend EP learning to both discrete and continuous complex-valued wave systems. In contrast to previous EP implementations, our scheme is valid in the weakly dissipative regime, and readily applicable to a wide range of physical settings, even without well defined nodes, where trainable inter-node connections can be replaced by trainable local potential."

The NEP protocol is defined for a system described by a time-dependent complex field ψ(r⃗, t). Training proceeds via a two-step procedure for each input sample. In the first, free-evolution phase, the system follows ∂t ψ(r⃗, t) = κψ(ψ(r⃗, t), θ, X), where θ is the vector of trainable parameters (such as local potentials, inter-node couplings, or input weights). The system is assumed to reach an oscillating steady state ψ(r⃗, t) = Ψ0(r⃗) e iωt, with frequency ω locked to a resonant drive ω = ωD. Moving to a rotating frame gives ∂t Ψ(r⃗, t) = κ(Ψ(r⃗, t), θ, X) = 0 in the steady state. In the second, nudged phase, a driving term proportional to the cost gradient is added: ∂t Ψ = κ − iβ ∂C(Ψ0, Ψ0)/∂Ψ0, where β is a small real parameter and the derivative is a Wirtinger derivative. The system relaxes to a new nudged steady state. The paper derives a local update rule for parameters: ∆θ ∝ −∂C/∂θ β=0 = i⟨∂Ψ/∂β, ∂κ/∂θ⟩ − i⟨∂Ψ/∂β, ∂κ/∂θ⟩, where ⟨f, g⟩ is the inner product. This rule generalizes previous work and holds for any complex wave equation, provided the near-equilibrium conditions ∂κ(r⃗)/∂Ψ(s⃗) = ∂κ(s⃗)/∂Ψ(r⃗) and ∂κ(r⃗)/∂Ψ(s⃗) ≈ −∂κ(s⃗)/∂Ψ(r⃗) are satisfied. The paper notes that "When a Hamiltonian functional H(Ψ, Ψ) exists such that κ ≈ i Ψ, H = i ∂H/∂Ψ with · as the Poisson bracket, Eq. (9) is naturally satisfied. As shown below, even when not exactly fulfilled, training can still converge successfully."

The framework is tested in exciton-polariton systems, described by a generalized continuous Gross–Pitaevskii equation (GPE) with dissipation γ, complex potential V(r⃗), resonant pumping P(r⃗), and nonlinear function f(Ψ) (either density response fd(Ψ) = gdΨ2 or saturation response fs(Ψ) = gs/(1+Ψ2)). The evolution equation is: ∂t Ψ ≡ κ = −(i/ħ)[−(ħ2/2m)∇2 + V + f(Ψ) − iγ]Ψ + P, where P(r⃗) = Σ k w k X k G(r⃗ − r⃗ k) with complex input weights w k. The trainable parameters are the complex potential V(r⃗) and pumping weights w k. The update rules derived are: for the potential, ∆V(r⃗) ∝ −∂C/∂V(r⃗) = −(1/ħ) ∂Ψ0(r⃗)2/∂β; for the pump weights, ∆w k ∝ −∂C/∂w k = −2∫ d⃗r X k G(r⃗ − r⃗ k) Im[∂Ψ0(r⃗)/∂β]. These updates depend on local steady-state fields obtained in the free and nudged phases. The paper emphasizes that The parameter β controls the nudging strength and must be small enough for accurate derivative estimation at β = 0, yet large enough to induce a measurable change in the wavefunction.

Numerical simulations use a discretized GPE on a finite spatial grid with Dirichlet boundary conditions, integrated with a fourth-order Runge–Kutta scheme. As a proof of principle, a one-dimensional polariton network with 9 nodes is trained to implement the two-input XOR function. Inputs are applied at sites 2 and 6, output is read at site 4. The cost function is the mean-squared error between output intensity and target intensity. With nudging parameter β = 0.01, learning rates lrV = lrw = 0.1, GPE parameters g = 0.1, γ = 0.1, and saturation nonlinearity, convergence occurs after about 10 epochs. Final outputs are Ψ42 = (0.00, 0.92, 1.06, 0.01) for inputs (00, 01, 10, 11). The paper reports: "Appendix B demonstrates that optimization remains effective regardless of input/output node positions and can rely on training both Vi and wi, or only Vi with fixed wi. The method is robust to random potential perturbations—e.g., structural inhomogeneities in experimental samples—verified by training with fixed small random fluctuations, which did not affect convergence."

For multi-class classification, a 2D discrete polariton network on a 15 × 150 grid is trained on the MNIST handwritten digits dataset (1000 samples per digit). The input region X corresponds to the entire network area, with each of ten 15 × 15 cells receiving a copy of the same input. The output region Y consists of ten central nodes, each representing one class. The cost function is categorical cross-entropy (CCE): CCCE(Ψ0, Ψ0) = −Σ i ΨY(r⃗ i)2 ln σ(Ψ0(r⃗ i)2), where σ is the softmax function. Convergence to a test accuracy of ≈ 90.3% is achieved after about 20 training epochs with β = 0.01, learning rates lrV = lrw = 0.1, GPE parameters g = 0.04, γ = 0.07, and density-response nonlinearity. The paper notes: Notably, the pumping weights shown in Fig. 4 form patterns that resemble handwritten digits, recognizable to the human eye. Varying the nonlinearity g reveals optimal performance for a moderate value g = 0.04, corresponding to a mean nonlinearity-to-loss ratio ⟨gΨ2⟩/(ħγ) ≈ 0.3. Larger values of g tend to destabilize learning. Training only the potential, without optimizing pumping weights, reduces performance but confirms purely optical optimization is possible. The paper compares: "A benchmark fully connected linear network with no hidden neurons and 10 outputs, trained by backpropagation, achieves ≈ 89.9% accuracy on the same MNIST test subset. Previous studies reported that oscillatory networks with local connections perform poorly on MNIST [23]. We attribute the improved performance to nonequilibrium wave dynamics, which overcome the locality of node couplings."

The paper provides an experimental outlook, stating: "The NEP protocol can be experimentally realized with current polariton systems. Numerical simulations for the 5-digit MNIST task (1000 samples per digit) converge after about 20 epochs for γ = 0.1, requiring two steady-state measurements per sample, each with ≈ 5000 RK4 steps and dt = 0.1. In physical units, the full training procedure completes within Tphys ≈ 0.1 ms to 10 ms for typical γ−1 ≈ 1 ps to 100 ps [35, 47], nearly six orders of magnitude faster than GPU-based numerical integration (NVIDIA GeForce RTX 2080 Ti, 11 GB)." For GaAs-based microcavities, typical parameters correspond to exciton density Ψ2 ≈ 30 µm−2, coupling energy of ≈ 1 meV, optically induced potential of ≈ 0.4 meV, and resonant pumping amplitude ≈ 0.8 meV, all within experimentally accessible range.

The paper concludes: "We have presented the NEP scheme, which transforms a nonlinear wave system into a machine-learning device trained via steady-state contrasts. Unlike previous approaches, our algorithm operates near equilibrium, trains local parameters instead of inter-node weights, and applies to both discrete and continuous systems. The protocol uses standard optical tools—SLMs and camera-based readouts—is robust to moderate inhomogeneities, and naturally supports tiling into weakly coupled cells for large-scale tasks. Altogether, these features make NEP a realistic framework for laboratory implementation, enabling in-situ, ultrafast, energy-efficient training on polaritonic hardware within experimental reach." The paper also notes that after submission, Scurria et al. reported an EP scheme for nonequilibrium steady states in real-valued state space, whereas NEP targets driven–dissipative polariton networks with complex fields and implements nudging via Wirtinger derivatives.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:

Improvement: Replace standard backpropagation with the Near-Equilibrium Propagation (NEP) algorithm for any AI system deployed on physical hardware (optical, photonic, or analog neuromorphic chips).

What the improved system can do:

  • Train directly on physical hardware without needing a digital model of the system, eliminating model–experiment mismatch

  • Perform in-situ training with only two steady-state measurements per sample (free and nudged phases), rather than requiring full computational graph traversal

  • Achieve comparable accuracy to backpropagation: 90.3% test accuracy on MNIST vs. 89.9% for a linear backprop network

  • Handle complex-valued wave dynamics (amplitude and phase) that standard EP cannot


Summary of capabilities: The improved AI system can train directly on physical wave-based hardware (photonic, polaritonic, or analog) with minimal measurements, handle complex-valued dynamics, tolerate dissipation and imperfections, train local parameters without explicit weights, and achieve accuracy comparable to digital backpropagation—all while operating orders of magnitude faster and more energy-efficiently than GPU-based training.

Sources

Related papers