Conservation laws determine what physical learning remembers
page_by_page
In short
The episode discusses a paper titled "Conservation laws determine what physical learning remembers." Hosts examine how physical rules like Equilibrium Propagation (EP), Coupled Learning (CL), and Adjoint Coupled Learning (AL) affect AI training memory. The discussion concludes that the conservation structure of learning rules dictates stability and inductive bias, with initial conditions having a permanent imprint on nonlinear systems.
Key concepts
- Conservation Laws
- These are physical rules applied to train circuits that determine how they learn. They dictate whether energy or mass is conserved during the training process, which fundamentally shapes what the AI system remembers about its initial state.
- Equilibrium Propagation (EP) and Coupled Learning (CL)
- These two rules are related and learn identical functions for single outputs. They conserve mass, meaning they imprint the initialization conditions onto the final result in a way that is stable across linear systems.
- Adjoint Coupled Learning (AL)
- This rule dissipates its conserved mass at a rate of twice its loss. This dissipation suggests a trade-off between training speed and performance, as it affects generalization quality in nonlinear circuits.
Terminology used across episodes
This episode discusses
- Conservation laws determine what physical learning remembers · Paper Radio
- Coercivity and Local Convergence of Physical Learning in Linear Circuits
- A Conservation Law for Equilibrium Propagation and Coupled Learning
The paper
Conservation laws determine what physical learning remembers · Read on arXiv
Bijaya Dangol
Independent Researcher
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Conservation laws determine what physical learning remembers".
Jane: The paper was written by Bijaya Dangol from Independent Researcher.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper summary: Tom: Welcome back to our show, "The Physics of AI," where we break down complex research into something digestible. We are talking about this fascinating paper, "Conservation laws determine what physical learning remembers." It’s a huge topic because it challenges the idea that all AI training is just about finding the optimal mathematical minimum.
Jane: Right, and this paper shows that when you use certain physical rules to train circuits—like those resistive networks—the way they are designed fundamentally dictates how they learn. It's not just about getting low error; it's about what kind of memory the training preserves.
Lu: This is profoundly important because the authors prove that in linear systems, these conservation laws are exact, which means the initialization you start with stays imprinted on the final result. That’s a huge theoretical shift from how we usually think about deep learning dynamics.
Meng: From an engineering perspective, this is big news because it suggests that when building hardware for AI, the physical design choices—the topology and those specific training rules—are not just constraints; they are defining the memory of the machine itself.
Lalam: I find it incredibly hopeful because if a learning device can permanently remember its initial conditions, it could lead to systems that are more stable and less prone to forgetting their original state, which has huge implications for reliability in AI applications.
Page 1 of the paper: Tom: So, let's look at the introduction on page one. The authors establish a context with three main rules: Equilibrium Propagation (EP), Coupled Learning (CL), and Adjoint Coupled Learning (AL). They show that EP and CL are related in a very specific way.
Jane: They prove that for a single output, these two rules learn identical functions up to just being reparameterized in time. That means if you only look at one result, you can't tell them apart based on the learned function itself.
Lu: But the authors point out that this equivalence breaks down as soon looking at multiple outputs, which is where the real differences start appearing. It’s about moving beyond simple demonstrations and seeing how these physical rules handle complexity.
Meng: And they introduce AL, Adjoint Coupled Learning, which doesn're very different because it doesn't conserve the mass at all; it actually dissipates it. This suggests that for a complex network with multiple outputs, the choice of rule becomes critical.
Lalam: It’s interesting to see this distinction, because if we are building physical hardware, choosing a path that preserves energy or choosing one that burns energy is a design choice with real consequences for stability.
Page 2 of the paper: Tom: Moving to page two, the authors set up their mathematical framework. They define the circuit as a graph with M trainable edges and Q outputs, but they assume that there are more edges than outputs, meaning M>O.
Jane: This condition is key because it means the solutions don't have a single fixed answer; instead, they form a manifold of possible solutions. The training process just moves along this surface until we reach a specific point on it.
Lu: The paper focuses on where that training ends up landing on that manifold, which determines the final learned function. This is different from standard AI optimization where we often just focus on convergence to a single best point.
Meng: In practical terms, since M>O, the configuration space is very rich. An engineer needs to know exactly how these constraints define the search space before they start running simulations on actual physical hardware.
Lalam: And we can see that by keeping track of where it lands, we are tracking not just an answer, but a unique historical imprint left by the initial conditions of the system.
Page 3 of the paper: Tom: On page three, they give us some powerful mathematical statements about these rules—Propositions one through to three. Proposition one confirms that EP and CL are trajectory equivalent for single outputs, which we already touched on.
Jane: But Proposition two is where AL stands out dramatically; the authors show it dissipates its own conserved mass at a rate of exactly twice its loss, meaning energy isn't preserved in that rule.
Lu: And Proposition three tells us why linearity is so special: in linear circuits, all three rules are homogeneous functions of degree-one zero and +one respectively. This means the scale doesn't change the direction of learning at all.
Meng: That homogeneity is a massive simplification for an engineer; it suggests that if you have a purely linear system, the initial scale doesn' not matter for how the circuit learns its function.
Lalam: It’s comforting to think that in certain systems, our initial setup is irrelevant to the final result, providing a stable starting point for cultural and technological applications.
Page 4 of the paper: Tom: The results on page four explore solution selection in linear circuits, where we expect that scale invariance. Table I shows that when we start with different initial conditions—different scaling—the learned input-output maps are incredibly close.
Jane: We see distances often in the ten-nine range, which is practically zero. This confirms Proposition three perfectly; the initialization memory is present but completely inert in linear systems.
Lu: The paper’s finding here is that for linear circuits, the inductive bias of physical learning is almost entirely dictated by where training begins, and the specific rules only provide a tiny correction to that bias.
Meng: This tells us that if we are building a large-scale linear AI chip, we don't need to worry about which rule—EP or CL—we pick in terms of function selection; they will all converge on the same result.
Lalam: It’s a powerful message that when systems are well-behaved and linear, their predictability is absolute, offering a sense of grounded reliability for complex AI systems.
Page 5 of the paper: Tom: Now, page five shifts gears by looking at nonlinear circuits where fixed elements like diodes are present. This is where the magic happens because Proposition three no longer holds.
Jane: We see that in these nonlinear setups, changing the initial scale factor—even if we keep the direction and task the same—can change the learned function by up to thirty-eight percent.
Lu: The nonlinearity allows that initialization memory to become functional, meaning the physical setup is imprinted on what it learns. The system remembers how conductive it was when it started training.
Meng: This is a critical takeaway for me, because if we use real-world components with non-linear characteristics, we must treat the initial state as a design parameter that significantly affects the final outcome.
Lalam: It’s a reminder that sometimes the most complex and unpredictable systems are also the ones that have the most memory, forcing us to be more deliberate about how we build them.
Page 6 of the paper: Tom: On page six, we look at generalization at matched training loss. The authors compare EP and CL against AL, Adjoint Coupled Learning. They find that while AL reaches the target faster, its generalization quality suffers.
Jane: At matched loss levels in nonlinear circuits, AL loses on average compared to EP and CLs. This is because of the mass dissipation that rule experiences along its flow path.
Lu: The fact that this penalty correlates with the amount of mass dissipated suggests a deep trade-off between speed and performance, which is a fundamental tension in designing any learning mechanism.
Meng: I'm interested in the data here, where they measure this "penalty." It shows that for larger circuits, where less mass is lost, AL’s penalty shrinks significantly.
Lalam: This trade-off highlights a balance we must strike: do we prioritize speed and efficiency in training, even at the cost of some generalization ability? We are constantly weighing performance against inherent stability.
Conclusion: Tom: To wrap up the discussion of "Conservation laws determine what physical learning remembers," we have seen how these three rules behave differently based on whether we are in a purely linear environment or a complex, nonlinear one.
Jane: We’ve established that while EP and CL conserve mass and retain memory, AL dissipates that mass, which translates into different training speeds and generalization capabilities.
Lu: The core message is that the conservation structure of these learning rules sets not just their stability but their "inductive bias," making the choice of a design parameter critical for what the machine will eventually compute.
Meng: This work gives us practical insights into designing physical AI hardware, showing that initial conditions are permanent imprints in nonlinear systems, which is something we can't ignore anymore.
Lalam: It’s a powerful message about the interplay between physics and computation; that our design choices aren’t just technical steps but decisions that have a lasting impact on the function and reliability of a learning system.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language