A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing

arXiv:2607.16183 · cs.LG, cs.ET, physics.app-ph · Submitted 2026-08-14 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing".

Jane: The paper was written by Owen Lockwood, Jérémy Béjanin, Joost Bus, Christopher Chamberland, Patrick Huembeli et al. from Extropic Corporation and Noumenal Labs Inc and University of Waterloo.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we're diving into a paper that's got a real mouthful of a title: "A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing."

Jane: And honestly, Tom, that title is doing a lot of heavy lifting. It's from the folks at Extropic Corporation, and the core idea is that instead of trying to fight the physics of heat and randomness in our computers, we should just embrace it.

Tom: Exactly! For decades, we've been building computers that try to be perfectly deterministic. We pump energy in to keep those electrons in line, to stop them from jiggling around. But this paper says, what if the jiggling *is* the computation?

Jane: Right. They're building what they call "thermodynamic neurons." These are physical systems, and in their case, superconducting circuits, that naturally settle into a state of equilibrium. And that equilibrium state, the probability of where a particle ends up, *is* the output of the computation.

Tom: So it's not a zero or a one. It's a distribution. It's like flipping a weighted coin, but the weight of that coin is something we can program.

Jane: And that's the "energy-based model" part. In machine learning, we have these things called Energy-Based Models, or EBMs, which are great in theory but a nightmare to run on normal digital hardware because you have to simulate all this randomness.

Tom: Right, it's computationally brutal. But if your hardware is *naturally* random, then sampling from that model is just... waiting for the physics to happen.

Jane: It's a beautiful inversion of the problem. We spend so much effort making our computers not-random, and then we spend even more effort simulating randomness in software. This paper is saying, let's just build the computer out of the randomness itself.

Tom: And the "continuous-variable" part? That's about the physical state, like the magnetic flux in a superconducting loop, being a continuous value, not a discrete bit.

Jane: Which gives you a lot more expressive power per physical element. You're not just encoding a binary state; you're encoding a whole range of possibilities.

Tom: So, the title is a promise. It's a blueprint for a whole new kind of computer. And we're going to spend the rest of the show figuring out if it delivers.

Jane: It's a big claim, but the physics is sound. The question is whether they can actually build it at scale.

Tom: And that's exactly what we're going to dig into next. Stick around.

Summary: Tom: So we've got this blueprint, and it's all about using physical systems that follow something called Langevin dynamics. Jane, can you break that down for our listeners?

Jane: Sure, Tom. Imagine you drop a marble into a bowl. It rolls around, and eventually, it settles at the bottom. That's a simple physical system finding its lowest energy state. Now, imagine the bowl isn't smooth, but has two dips in it. The marble will end up in one of them, but which one is random, depending on how it was dropped.

Tom: And the "Langevin dynamics" is the math that describes that rolling, bouncing, and settling process.

Jane: Exactly. It includes the forces from the shape of the bowl, the friction that slows the marble down, and the random thermal kicks from the environment. The paper shows that if you can build a physical system with a programmable "bowl" shape, the probability of finding the marble in any given spot follows a very specific, very useful distribution called the Gibbs distribution.

Tom: And that Gibbs distribution is the backbone of those energy-based models we mentioned. So by controlling the shape of the potential, you control the probability distribution of the output.

Jane: Right. And they show how to build basic "bowls" — a single well for a Gaussian distribution, a double well for a binary-like distribution, and then, crucially, how to couple these systems together.

Tom: Coupling is where it gets interesting. That's how you build complex models. They show that by coupling two of these "neurons" together, you can compute things like a sigmoid function, which is a core building block in neural networks.

Jane: And they don't stop there. They outline how to build a matrix-vector product, a softmax function, and even a full transformer architecture. They call it the "Thermoformer."

Tom: A thermodynamic transformer. That's wild. And the key insight is that you don't just get a single sample from these systems. You can measure the *average* state over time, which gives you a very precise, deterministic-looking output.

Jane: And that's how they bridge the gap between probabilistic physics and the deterministic math we use in modern AI. It's a way to do the same calculations, but the heavy lifting is done by the laws of thermodynamics, not by a clock and a bunch of transistors.

Tom: So the summary is: they've figured out a way to map the core operations of machine learning onto physical systems that naturally perform those operations just by existing. It's a pretty elegant idea.

Jane: It is. And the potential payoff, which we'll get into, is a massive reduction in energy consumption and a potential speedup. But first, let's talk about how they actually plan to build this thing.

Improvements: Tom: So the blueprint is there, but what's the actual hardware? Jane, you mentioned superconducting circuits. What does that look like?

Jane: So, they're using something called a Josephson junction. It's a tiny, non-linear electronic component that only works at very low temperatures. It acts like a tunable inductor, and it allows them to create that double-well potential we talked about.

Tom: And they actually built one! They fabricated a chip with these "thermodynamic neurons" and ran experiments. What did they find?

Jane: They showed they could control the barrier between the two wells, and they measured the escape rate of the "flux particle" from one well to the other. This is a direct measurement of how the system thermalizes.

Tom: And that thermalization time is the speed of the computation. They showed that at low temperatures, the escape is dominated by quantum tunneling, and at higher temperatures, it's thermally activated, which matches their theoretical model.

Meng: But from an engineering standpoint, I have to ask about the elephant in the room. This is a single, isolated neuron. The paper talks about coupling thousands of these together to do something useful. How do you wire that up?

Jane: That's the big challenge, Meng. They mention tunable inductive couplers, similar to what's used in quantum annealers. But they also acknowledge that all-to-all connectivity is impossible, so you have to work with a sparse graph and use latent variables to mediate interactions.

Meng: And what about the readout? You can't just hook a wire up to a superconducting circuit without disturbing it.

Jane: They use a dispersive readout, which is a standard technique in superconducting qubits. It's a bit destructive, as they say, because it interrupts the state to measure it. But they can get a binary "left well" or "right well" answer.

Tom: And for the more complex operations, like computing an average, they propose something called "estimation oscillators." It's a clever idea where you couple the neuron you want to measure to a bunch of other neurons, let them equilibrate, and then the average of *their* states gives you the answer.

Meng: So it's a way to do analog averaging, which is a huge deal. It avoids the need for a fast, high-precision analog-to-digital converter, which would be a bottleneck.

Jane: Exactly. It's a full-stack approach. They're not just proposing a new transistor; they're proposing a new way to think about the entire computing pipeline, from the physics up to the software.

Tom: So the improvements aren't just incremental. It's a new paradigm. But what does that mean for the world? Let's bring in Lu and Lalam to get their take on the big picture.

Conclusion: Tom: So, Lu, you've been listening to all this. What's the big, world-changing implication of "A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing"?

Lu: Tom, this is a potential answer to the energy wall we're hitting with AI. We're building massive data centers that consume enormous amounts of power. This paper suggests a path where the computation itself is done at the thermodynamic limit, using only the ambient thermal noise as the engine.

Jane: And it's not just about energy. The paper argues that these systems could be fundamentally faster for sampling-based tasks, which are everywhere in modern AI, from generative models to reinforcement learning.

Meng: But I'm still skeptical about the practical timeline. The paper is a blueprint, but we're talking about a single, hand-tuned device in a dilution fridge. Scaling that up to a useful machine is a decade-long engineering problem.

Lalam: I agree, Meng, but consider the cultural impact. This isn't just a faster computer. It's a fundamental shift in how we perceive computation. It moves us away from the rigid, deterministic logic of the past and towards a more probabilistic, physics-based understanding.

Tom: That's a beautiful way to put it, Lalam. It's like we're learning to speak the language of the universe, which is inherently probabilistic at its core.

Lalam: And that could democratize AI. If the hardware is more energy-efficient, it lowers the barrier to entry. Smaller companies, research labs, even individuals could train and run models that are currently only possible for tech giants with massive power infrastructure.

Jane: So, to wrap it up, this paper gives us a concrete, physics-grounded roadmap for a new kind of computer. It's a huge leap, but the foundational experiments are already showing promise.

Tom: And while it might be a while before we see a "Thermoformer" powering our phones, the ideas here are going to shape the next generation of computing. It's a fantastic paper to get our heads around.

Jane: Absolutely. We've covered the theory, the hardware, and the potential. It's a lot to digest, but it's an exciting direction.

Tom: Thanks for joining us on this deep dive. We'll be back next time to break down another paper from the arXiv. Until then, keep questioning the bits.

Jane: And maybe, just maybe, embrace the noise. See you all next time.

Owen Lockwood, Jérémy Béjanin, Joost Bus, Christopher Chamberland, Patrick Huembeli, Frank Schäfer, Guillaume Verdon

Extropic Corporation · Noumenal Labs Inc · University of Waterloo

cs.LG, cs.ET, physics.app-ph

Submitted: 2026-08-14

Updated: 2026-08-18

Comments: 42 pages, 20 figures

Code: https://github.com/lockwo/distreqx

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 74/100

The gist: The paper introduces a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware.

Key concepts

Thermodynamic Neurons
These are physical systems, such as superconducting circuits, designed to naturally settle into a state of equilibrium. The probability of where a particle ends up in this stable state is the output of the computation, moving away from traditional binary logic.
Langevin Dynamics
This is the mathematical description that governs how a physical system settles into its lowest energy state. It accounts for forces, friction, and random thermal kicks from the environment as a particle moves through a potential 'bowl' shape.
Energy-Based Models (EBMs)
These models are based on the probability distribution of a physical system at equilibrium. By controlling the shape of the physical potential (the 'bowl'), researchers can program and control exactly what that resulting probability distribution will be.
The Thermoformer
This is a thermodynamic version of complex neural network architectures, like transformers. It is built by coupling multiple thermodynamic neurons together, allowing them to compute core functions such as sigmoid and softmax using physical interactions.

Terminology

Summary

The paper introduces a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. The authors state: "we demonstrate a thermodynamic computing paradigm based on the equilibria of energy functions which undergo Langevin dynamics, and which has potential to impact the computing landscape through increased time and energy efficiency. Thermodynamic computing as a method is distinguished from other kinds of computing in that it harnesses stochastic fluctuations as a core computing resource."

The core concept is that "the steady-state distribution of physical Langevin dynamics is a Gibbs distribution with energy potential Uθ(x). If we can design a physical system with a sufficiently controllable energy potential landscape, so that the time and energy consumption for physical equilibration and readout is smaller than for digital MCMC sampling, we expect computational benefits from such a thermodynamic computing device."

The paper contrasts this with classical and quantum computing: "In classical (that is, digital and deterministic) and quantum computing stacks, much effort is put into the elimination of random fluctuations from within the system... Yet in modern algorithms, we often are forced to reintroduce stochastic fluctuations (as these algorithms rely on noisy gradients, Monte Carlo estimates, sampling-based inference, etc.), despite having engineered them out of the hardware to the best of our ability."

The authors present a framework based on a set of M thermodynamic neurons, which are composable subsystems S, each described by Langevin dynamics [12, 13], such that the full state space is X = S1 × ⋅ ⋅ ⋅ × SM.

Key theoretical foundations include the Gibbs distribution: πθ (x) = 1/Zθ e−βEθ (x) which has inspired a class of machine learning models called energy-based models (EBMs). The paper notes that "EBMs are highly flexible machine learning ansätze that have potent features like composability and the ability to do conditional inference through clamping. However, EBMs are difficult to scale on traditional digital hardware due to the computational cost of sampling routines dealing with the intractable normalization constant Zθ."

The paper details the Langevin dynamics approach: "The celebrated Langevin Monte Carlo method [60] represents such an approach to draw samples (x ∈ RD) from an EBM... samples can be obtained from the path of the Langevin diffusion process, whose states x(t) have distribution P (t) and evolve according to the stochastic ordinary differential equation γi dxi = − ∂Uθ (x)/∂xi dt + √(2γi/β) dWt(i)."

For physical implementations, the authors use underdamped Langevin dynamics: dxi = pi/mi dt, dpi = −(∂Uθ (x)/∂xi + γi/mi pi) dt + √(2γi/β) dWt(i) and note that the Fokker-Planck equation associated with underdamped Langevin dynamics has the equilibrium solution πθ (x, p) = e−βEθ (x,p)/Zθ.

The paper discusses energy-time-precision trade-offs: "The Monte Carlo standard error scales as O(N−1/2)... A key advantage of a thermodynamic approach is that precision is independent of the magnitude of computed values... The number of samples N can be adjusted to meet the required precision for each computation."

Regarding Landauer's principle: "Landauer's principle states that erasing a bit of information must dissipate at least kB T ln(2) energy... This limit is many orders of magnitude lower than what is currently used in deterministic digital computing systems... computation based on thermodynamic principles could potentially operate with energy consumption much closer to Landauer's limit."

The paper introduces elemental potentials: "The two main classes of single-particle potentials we focus on are Gaussian (also called single-well or quadratic), where the force is affine in x (here, x ∈ R), and nonlinear (also called quartic or double-well) potentials. The Gaussian potential is Uθsw (x) = 1/2σ (x − µ)2 and the double-well is Uθdw (x) = λ1 x2 (x − 1)2 − λ2 x."

Coupling potentials include: Uθsig (x, z) = λ1 x2 (x − 1)2 − zx + z2 which allows us to program the commonly used sigmoid activation σML: f (z) = lim λ1→∞ Ex∼πθ (⋅∣z) [x] = 1/(1 + exp(−z)) = σML (z).

Other coupling potentials enable matrix-vector products: UθMVP (x, z) = ν/2 ∑i x2i − x⊺ W z which results in the expected value of x being the matrix vector product (MVP) W z.

The paper demonstrates how to construct deep learning models: By working with expectations rather than individual samples, we can recover (almost) deterministic operations similar to common digital subroutines while potentially consuming less energy. It shows how to implement MLPs, softmax (via Uθsoft (x, z) = λ1 ∑i x2i (xi − 1)2 − ∑i xi zi + λ2 (∑i xi − 1)2), and transformers (thermoformer).

For gradient computation, the paper derives: ∂ E[y]/∂θ = −Cov (y, ∂Eθ (y∣x)/∂θ) and ∂ E[y]/∂x = −Cov (y, ∂Eθ (y∣x)/∂x). These Jacobians can be used to forward- and back-propagate derivative information.

The paper introduces probabilistic graphical models: Using the factor graph formalism, we can represent both directed and undirected graphical models [130], where models are expressed as bipartite graphs G = (Vv, Vf, E). This allows π(xVv) = 1/Z ∏a∈Vf exp [−Ea (xn(a))].

Example architectures include Gaussian PGMs, Gaussian mixture models, hidden Markov models, continuous Ising models, and transformers. The paper notes: Even when using only simple Gaussian potentials, powerful models can be constructed.

For on-chip self-learning, the paper proposes: a timescale-separated Langevin system in which the fast variables (x, z) rapidly equilibrate for quasi-static θ, while θ drifts under an effective force that encodes the learning signal. The effective Born-Oppenheimer force is FiBO (θ) ∶= − ∫ ddz z ddx x e−U (x,z,θ)/Z(θ) ∂U (x, z, θ)/∂θi.

The experimental implementation uses superconducting circuits: We designed and fabricated a superconducting chip that implements the core double-well building block, a thermodynamic neuron [205]. The device comprises a CPW section similar to that of a λ/4 resonator but shunted to ground at the open end by a dc-SQUID loop formed by a pair of Josephson junctions.

The Hamiltonian is H = 1/2C q2 + 1/2L (ϕ − ϕtilt)2 − EJ (ϕbar) cos (2π ϕ/Φ0) where the charge q and flux ϕ are conjugate variables.

The paper describes the experimental results: We plot population P data from the relaxation experiment for a few select barrier settings and temperatures... the average population decays exponentially until it stabilizes at approximately P = 0.5. The escape energy Eesc as a function of T shows At low temperature, the escape energy appears to be constant. Above 100 mK, Eesc increases linearly until approximately 165 to 200 mK, at which point the escape energy stops increasing and instead appears to taper off.

The paper concludes: "We presented a framework for energy-based thermodynamic computing based on probabilistic graphical models and superconducting circuits driven by thermal fluctuations. We have described the relevant building blocks and carried out idealized analyses of equilibration times for relevant operations of the device. These results suggest a variety of avenues for further explorations."

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Replace digital Langevin Monte Carlo sampling with physical Langevin dynamics on superconducting hardware. The paper demonstrates that thermal noise in a tunable double-well potential naturally samples from a Gibbs distribution, eliminating the need for iterative digital sampling.

What the improved AI system can do:

  • Train EBMs using contrastive divergence without the computational bottleneck of negative-phase sampling

  • Sample from high-dimensional distributions in nanoseconds instead of milliseconds

  • Achieve sampling rates limited only by physical thermalization time (τ therm), not digital iteration count


Key limitation to note: These improvements require superconducting hardware operating at millikelvin temperatures, which is not yet commercially available at scale. However, the paper demonstrates working prototypes and provides a clear path toward integration with existing AI infrastructure.

Abstract

To address the escalating energy and latency demands of machine-learning workloads, we introduce a blueprint for an energy-efficient and fast thermodynamic computing stack that leverages stochastic analog processes in physical hardware. In this work, we focus on energy-based thermodynamic computing where the stochastic process is well described by Langevin dynamics with tunable energy potentials. The implementation of such potentials in physical hardware enables us to generate and sample from basic parameterized energy-based models. We demonstrate how to construct and train popular classes of machine learning models based on these hardware-native energy-based models, using the framework of probabilistic graphical models. We analyze the runtime and energy consumption of different models in this thermodynamic paradigm based on theoretical considerations and numerical studies. As a preliminary experimental realization of such hardware, we present our stochastic analog superconducting circuits driven by thermal noise. Together, these results outline a path toward energy-efficient thermodynamic hardware for probabilistic machine learning.

Sources

Related papers