Conservation laws determine what physical learning remembers

arXiv:2608.00097 · cond-mat.soft, cond-mat.dis-nn, cs.AI, cs.LG · Submitted 2026-08-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Conservation laws determine what physical learning remembers".

Jane: The paper was written by Bijaya Dangol from Independent Researcher.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Paper summary: Tom: Welcome back to our show, "The Physics of AI," where we break down complex research into something digestible. We are talking about this fascinating paper, "Conservation laws determine what physical learning remembers." It’s a huge topic because it challenges the idea that all AI training is just about finding the optimal mathematical minimum.

Jane: Right, and this paper shows that when you use certain physical rules to train circuits—like those resistive networks—the way they are designed fundamentally dictates how they learn. It's not just about getting low error; it's about what kind of memory the training preserves.

Lu: This is profoundly important because the authors prove that in linear systems, these conservation laws are exact, which means the initialization you start with stays imprinted on the final result. That’s a huge theoretical shift from how we usually think about deep learning dynamics.

Meng: From an engineering perspective, this is big news because it suggests that when building hardware for AI, the physical design choices—the topology and those specific training rules—are not just constraints; they are defining the memory of the machine itself.

Lalam: I find it incredibly hopeful because if a learning device can permanently remember its initial conditions, it could lead to systems that are more stable and less prone to forgetting their original state, which has huge implications for reliability in AI applications.

Page 1 of the paper: Tom: So, let's look at the introduction on page one. The authors establish a context with three main rules: Equilibrium Propagation (EP), Coupled Learning (CL), and Adjoint Coupled Learning (AL). They show that EP and CL are related in a very specific way.

Jane: They prove that for a single output, these two rules learn identical functions up to just being reparameterized in time. That means if you only look at one result, you can't tell them apart based on the learned function itself.

Lu: But the authors point out that this equivalence breaks down as soon looking at multiple outputs, which is where the real differences start appearing. It’s about moving beyond simple demonstrations and seeing how these physical rules handle complexity.

Meng: And they introduce AL, Adjoint Coupled Learning, which doesn're very different because it doesn't conserve the mass at all; it actually dissipates it. This suggests that for a complex network with multiple outputs, the choice of rule becomes critical.

Lalam: It’s interesting to see this distinction, because if we are building physical hardware, choosing a path that preserves energy or choosing one that burns energy is a design choice with real consequences for stability.

Page 2 of the paper: Tom: Moving to page two, the authors set up their mathematical framework. They define the circuit as a graph with M trainable edges and Q outputs, but they assume that there are more edges than outputs, meaning M>O.

Jane: This condition is key because it means the solutions don't have a single fixed answer; instead, they form a manifold of possible solutions. The training process just moves along this surface until we reach a specific point on it.

Lu: The paper focuses on where that training ends up landing on that manifold, which determines the final learned function. This is different from standard AI optimization where we often just focus on convergence to a single best point.

Meng: In practical terms, since M>O, the configuration space is very rich. An engineer needs to know exactly how these constraints define the search space before they start running simulations on actual physical hardware.

Lalam: And we can see that by keeping track of where it lands, we are tracking not just an answer, but a unique historical imprint left by the initial conditions of the system.

Page 3 of the paper: Tom: On page three, they give us some powerful mathematical statements about these rules—Propositions one through to three. Proposition one confirms that EP and CL are trajectory equivalent for single outputs, which we already touched on.

Jane: But Proposition two is where AL stands out dramatically; the authors show it dissipates its own conserved mass at a rate of exactly twice its loss, meaning energy isn't preserved in that rule.

Lu: And Proposition three tells us why linearity is so special: in linear circuits, all three rules are homogeneous functions of degree-one zero and +one respectively. This means the scale doesn't change the direction of learning at all.

Meng: That homogeneity is a massive simplification for an engineer; it suggests that if you have a purely linear system, the initial scale doesn' not matter for how the circuit learns its function.

Lalam: It’s comforting to think that in certain systems, our initial setup is irrelevant to the final result, providing a stable starting point for cultural and technological applications.

Page 4 of the paper: Tom: The results on page four explore solution selection in linear circuits, where we expect that scale invariance. Table I shows that when we start with different initial conditions—different scaling—the learned input-output maps are incredibly close.

Jane: We see distances often in the ten-nine range, which is practically zero. This confirms Proposition three perfectly; the initialization memory is present but completely inert in linear systems.

Lu: The paper’s finding here is that for linear circuits, the inductive bias of physical learning is almost entirely dictated by where training begins, and the specific rules only provide a tiny correction to that bias.

Meng: This tells us that if we are building a large-scale linear AI chip, we don't need to worry about which rule—EP or CL—we pick in terms of function selection; they will all converge on the same result.

Lalam: It’s a powerful message that when systems are well-behaved and linear, their predictability is absolute, offering a sense of grounded reliability for complex AI systems.

Page 5 of the paper: Tom: Now, page five shifts gears by looking at nonlinear circuits where fixed elements like diodes are present. This is where the magic happens because Proposition three no longer holds.

Jane: We see that in these nonlinear setups, changing the initial scale factor—even if we keep the direction and task the same—can change the learned function by up to thirty-eight percent.

Lu: The nonlinearity allows that initialization memory to become functional, meaning the physical setup is imprinted on what it learns. The system remembers how conductive it was when it started training.

Meng: This is a critical takeaway for me, because if we use real-world components with non-linear characteristics, we must treat the initial state as a design parameter that significantly affects the final outcome.

Lalam: It’s a reminder that sometimes the most complex and unpredictable systems are also the ones that have the most memory, forcing us to be more deliberate about how we build them.

Page 6 of the paper: Tom: On page six, we look at generalization at matched training loss. The authors compare EP and CL against AL, Adjoint Coupled Learning. They find that while AL reaches the target faster, its generalization quality suffers.

Jane: At matched loss levels in nonlinear circuits, AL loses on average compared to EP and CLs. This is because of the mass dissipation that rule experiences along its flow path.

Lu: The fact that this penalty correlates with the amount of mass dissipated suggests a deep trade-off between speed and performance, which is a fundamental tension in designing any learning mechanism.

Meng: I'm interested in the data here, where they measure this "penalty." It shows that for larger circuits, where less mass is lost, AL’s penalty shrinks significantly.

Lalam: This trade-off highlights a balance we must strike: do we prioritize speed and efficiency in training, even at the cost of some generalization ability? We are constantly weighing performance against inherent stability.

Conclusion: Tom: To wrap up the discussion of "Conservation laws determine what physical learning remembers," we have seen how these three rules behave differently based on whether we are in a purely linear environment or a complex, nonlinear one.

Jane: We’ve established that while EP and CL conserve mass and retain memory, AL dissipates that mass, which translates into different training speeds and generalization capabilities.

Lu: The core message is that the conservation structure of these learning rules sets not just their stability but their "inductive bias," making the choice of a design parameter critical for what the machine will eventually compute.

Meng: This work gives us practical insights into designing physical AI hardware, showing that initial conditions are permanent imprints in nonlinear systems, which is something we can't ignore anymore.

Lalam: It’s a powerful message about the interplay between physics and computation; that our design choices aren’t just technical steps but decisions that have a lasting impact on the function and reliability of a learning system.

Bijaya Dangol

Independent Researcher

cond-mat.soft, cond-mat.dis-nn, cs.AI, cs.LG

Submitted: 2026-08-19

Updated: 2026-08-20

Comments: 8 pages, 3 figures, 3 tables

Code: https://github.com/dangoldbj/physical-learning-memory

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 74/100

Key concepts

Conservation Laws
These are physical rules applied to train circuits that determine how they learn. They dictate whether energy or mass is conserved during the training process, which fundamentally shapes what the AI system remembers about its initial state.
Equilibrium Propagation (EP) and Coupled Learning (CL)
These two rules are related and learn identical functions for single outputs. They conserve mass, meaning they imprint the initialization conditions onto the final result in a way that is stable across linear systems.
Adjoint Coupled Learning (AL)
This rule dissipates its conserved mass at a rate of twice its loss. This dissipation suggests a trade-off between training speed and performance, as it affects generalization quality in nonlinear circuits.

Terminology

Summary

Summary of Conservation laws determine what physical learning remembers

This paper investigates how local learning rules—specifically Equilibrium Propagation (EP), Coupled Learning (CL), and Adjoint Coupled Learning (AL)—train resistive networks, focusing on the role of conservation laws in determining the final learned input-output function.

Theoretical Framework and Core Propositions

The study utilizes a framework where a circuit is defined by N nodes and M trainable edges, each carrying a conductance kappa e. The training process involves minimizing co-content, leading to three distinct continuous-time flows: EP, CL, and AL.

  1. Indistinguishability in Single Outputs: For single output scenarios (O=1, where inputs and outputs are disjoint), the EP and CL flows have identical orbits up to a time reparametrization (Proposition 1). This means that single-output experiments... cannot reveal an inductive-bias difference between EP and CL.

  2. The Dissipative Nature of AL: Adjoint Coupled Learning (AL) is fundamentally different from the others. It does not conserve the mass K; instead, it dissipates it at a precise rate: = - mu 0 squared = -2* (Proposition 2), where* is the adjoint loss.

3 in Linear Circuits: In linear resistive circuits, the conserved mass (K) is functionally inert. The vector fields for EP, CL, and AL are homogeneous in kappa of degrees-1, 0, and +1, respectively (Proposition 3). This homogeneity ensures that the selected solution is independent of the initialization scale.

Experimental Validation: Linear Circuits (Where Theory Holds)

Experiments conducted on linear circuits confirm these theoretical constraints:

  • Selection Differences: When two outputs are used (O=2), EP and CL land on measurably different points of the solution manifold, with a median relative Frobenius distance (delta) of 1.8 times 10-5. AL lands roughly a hundred times farther from either.

  • Inductive Bias: In linear circuits, the inductive bias is set almost entirely by where training starts, meaning the initial conditions determine the outcome, and the rule contributes a small systematic correction. The scale-invariant nature of these systems means that in linear controls, initialization memory is inert.

Experimental Validation: Nonlinear Circuits (Where Theory Breaks Down)

The study then examines circuits containing fixed, untrained elements (like diodes), where nonlinearity becomes functional:

  • Functionality of Scale: In these nonlinear circuits, the scale effect becomes large and functional. Rescaling the initial conductance (kappa 0 by up to c=4 changes the learned function by up to thirty-eight percent.

  • The Dissipation Effect: The dissipative rule (AL) shows a reduced dependence on initialization scale compared to EP or CL. Its mass decay drives trajectories from different initial spheres toward one another and erases part of the scale memory that conservative rules are bound to keep.

  • Leaky Conservation: Even in linear circuits, when fixed elements are introduced, EP and CL conserve K only approximately (a median relative drift of 1.5 times 10-3).

Generalization at Matched Training Loss

The paper assesses the quality of the solutions selected by each rule:

  • At matched training loss, AL generalizes worse than EP and CL in a significant majority of paired comparisons. The penalty correlates with the mass dissipated en route and fades in larger circuits where little mass is lost.

  • The difference in generalization is attributed to the dissipation of K.

Conclusion

The conservation structure of a local learning rule dictates its behavior: it sets its initialization memory, its training speed, and... its generalization. A designer's choice among EP, CL, and AL is fundamentally a choice about how much the trained network will remember its fabrication state (conservation), how fast it will train (dissipation), and how well it will generalize.

Improvements for AI systems

Based on a detailed analysis of this scientific paper, I have identified three specific improvements that can be implemented in modern AI systems, along with the precise capabilities these improvements grant to the resulting architecture.

Improvement: Integrate a training mechanism that enforces exact conservation of a defined structural property (analogous to K in the paper) during learning, specifically using an EP or CL-like dynamic. This replaces traditional stochastic regularization methods with a deterministic, physics-based constraint.

What the Improved AI System Can Do:

  • Achieve Permanent Initialization Memory: The system will retain a permanent imprint of its initial weights/structure (initialization scale), which is not transient or lost through standard training. This allows the network to act as a highly deterministic, physics-informed lookup table for specific tasks.

  • Ensure Structural Stability: By constraining the learning trajectory to maintain this conserved mass, the system becomes inherently stable against drift and operational range violations, making it ideal for high-reliability hardware or environments where stability is paramount.

  • Maintain Scale Invariance (When Applicable): In simplified or linear contexts, this feature ensures that the learned function is independent of initial scale, providing predictable behavior across various input magnitudes.

Improvement: Implement a dynamic loss control mechanism that allows the AI system to switch between three defined training regimes:

  1. Conservative Mode (EP/CL): Prioritizing stability and memory retention over convergence speed.

  2. Dissipative Mode (AL): Prioritizing rapid convergence toward a solution, allowing controlled mass dissipation (= -2*).

  3. Hybrid/Adaptive Mode: A system that dynamically selects the optimal mode based on the required trade-off between training speed and generalization performance.

What the Improved AI System Can Do:

  • Optimize for Convergence Speed: When rapid training is prioritized (e.g., initial model deployment or resource-constrained environments), the dissipative AL regime allows the system to reach target loss levels significantly faster than conservative methods.

  • Ensure High Generalization Quality: Once a solution is found, or when generalization is critical, the system can revert to a conservative mode, ensuring that the selected function maintains high fidelity and minimal function-space error (delta).

  • Manage Trade-offs Explicitly: The AI system can be designed to perform trade-off analysis, choosing between the faster (but potentially less generalizable) AL approach and the slower but more robust EP/CL approach, based on the specific deployment requirements.

Improvement: Design architectures that deliberately incorporate fixed, non-trainable nonlinear elements (analogous to diodes or rectifiers) alongside trainable components. This leverages the functional dependence on initialization scale (c) inherent in nonlinear systems.

What the Improved AI System Can Do:

  • Implement Controlled Inductive Bias: The system utilizes the initialization memory as a controllable feature. By selecting a specific initial conductance scale (c), the network can be biased toward a known, high-performing functional state. This is not merely random initialization; it is functional initialization.

  • Achieve High Precision for Complex Tasks: Unlike linear systems where initialization memory is inert, this architecture allows the initial conditions to dictate a measurable shift in the learned function (up to 40% in analogous examples). This capability makes the system suitable for tasks requiring precise, non-trivial functional mappings that are not guaranteed by simple gradient descent.

  • Mitigate Washout Effects: The system can leverage this memory to maintain structural fidelity over long training periods or complex operations, ensuring that the initial design intent is preserved throughout its operational life.

Sources

Related papers