Differentiable Thermodynamic Phase-Equilibria for Machine Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Differentiable Thermodynamic Phase-Equilibria for Machine Learning".
Jane: The paper was written by Karim K. Ben Hicham, Moreno Ascani, Jan G. Rittig and Alexander Mitsos from RWTH Aachen University and Forschungszentrum Jülich GmbH and JARA-ENERGY.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Alright, welcome back to the show, everyone. We've got a paper that's got both of us genuinely excited, and it's called "Differentiable Thermodynamic Phase-Equilibria for Machine Learning." Jane, I'll be honest, when I first read that title, I had to read it twice.
Jane: You and me both, Tom. But once you unpack it, it's actually a beautiful idea. These researchers from RWTH Aachen and Forschungszentrum Jülich, led by Alexander Mitsos, they're trying to teach eye to understand a really fundamental physics problem: when you mix two liquids, do they separate into two layers, and what are the compositions of those layers?
Tom: Right, like oil and water. But they're not just predicting it, they're making sure the eye actually follows the laws of thermodynamics while it learns. That's the "differentiable" part, right?
Jane: Exactly. Normally, when you train a neural network to predict something, you just show it examples and let it figure out the pattern. But here, the thing they want to predict—the equilibrium state—is defined by a minimization problem. It's the state that has the lowest possible Gibbs energy.
Tom: And that minimization problem is a real pain in the neck because it's non-convex. There can be multiple local minima, and you need to find the global one. That's not something you can just differentiate through easily.
Jane: And that's the crux of it. If you can't differentiate through the solver, you can't train the network end-to-end. You're stuck. So what they did is they built a solver that works on a discrete grid of compositions, finds the true minimum in the forward pass, and then uses a clever trick called a straight-through gradient estimator to get gradients in the backward pass.
Tom: It's like they're saying, "We'll use the real physics to make the prediction, but we'll use a smooth, fuzzy version of that physics to figure out how to improve." And the results are pretty striking. On their test set, they got a mean absolute error of about zero point zero six eight, which is a solid improvement over the previous surrogate-based method that got zero point zero eight zero.
Jane: And the key thing is, their method is thermodynamically consistent by construction. The surrogate method, which trains a separate neural network to mimic the solver, doesn't guarantee that. It can give you answers that violate the laws of physics.
Tom: So we're not just getting better numbers, we're getting numbers that actually mean something physically. That's a huge deal for chemical engineering.
Jane: It really is. And I think the implications go way beyond just liquid-liquid equilibrium. This framework could be applied to vapor-liquid equilibria, solid-liquid equilibria, even more complex systems. It's a general way to embed physics into machine learning.
Tom: And that's exactly what we're going to dig into next. We're going to look at how they actually built this thing, the algorithm itself, and why it's such a clever piece of work.
Summary: Jane: So, Tom, we've established that this paper, "Differentiable Thermodynamic Phase-Equilibria for Machine Learning," is a big deal. But let's get into the nitty-gritty of how they actually pulled it off.
Tom: Please, because the summary in the abstract is dense. They mention "discrete enumeration of feasible phase states" and "masked softmax aggregation." I need that in plain English.
Jane: Okay, so imagine you have a binary mixture. You know the overall composition, let's call it z. The system can either be one single phase, or it can split into two phases with compositions x' and x''. The total Gibbs energy of the mixture depends on those two compositions.
Tom: And the equilibrium is the pair (x', x'') that gives the lowest total Gibbs energy. Got it.
Jane: Now, their algorithm, which they call DISCOMAX, doesn't try to solve that continuous optimization problem directly. Instead, it creates a grid of possible compositions, say one hundred one points between zero and one. Then it enumerates every possible pair of grid points that could satisfy the mass balance for the given feed z.
Tom: So it's a brute-force search over a discretized space. That's the "discrete enumeration" part.
Jane: Exactly. And for each of those pairs, it calculates the total Gibbs energy. Then, in the forward pass, it just picks the pair with the lowest energy. That's the hard, thermodynamically exact answer, up to the resolution of the grid.
Tom: But that hard minimum is a piecewise-constant function. The gradient is zero almost everywhere, which is useless for training.
Jane: Right, and that's where the "masked softmax" comes in. Instead of just picking the minimum, they compute a softmax distribution over all the feasible pairs. The lower the Gibbs energy of a pair, the higher its weight in this distribution. Then they compute a weighted average of the compositions.
Tom: So it's a smooth, differentiable approximation of the argmin.
Jane: Precisely. And then they use the straight-through estimator. The forward pass uses the hard minimum, so the prediction is physically correct. But the backward pass uses the gradients from the soft, weighted average. It's the best of both worlds.
Tom: And this softmax thing, it's not just a mathematical trick. It's actually analogous to the Boltzmann distribution in statistical thermodynamics. The probability of a system being in a certain state is proportional to the exponential of negative energy over temperature.
Jane: That's a beautiful connection. The temperature parameter τ in their softmax plays the same role as temperature in the Boltzmann distribution. As τ goes to zero, the distribution sharpens and concentrates on the true minimum. It's like they're using statistical mechanics to guide the learning.
Tom: And that's why the method is so elegant. It's not just a hack to get gradients; it's grounded in the physics of the problem.
Jane: Exactly. And the fact that they can train a graph neural network to predict the excess Gibbs energy, and then differentiate through this entire equilibrium calculation, is a real feat of engineering.
Tom: I'm curious about the practical side, though. How does this compare to just training a neural network to be a surrogate solver, like the previous work did?
Jane: Well, that's exactly what we're going to talk about next. The surrogate approach has some fundamental limitations, and this paper shows them pretty clearly.
Improvements: Tom: So, Jane, we've got the algorithm. Now, what does this paper actually improve upon? What was the state of the art before this?
Jane: The main baseline is a surrogate solver approach from a previous paper. The idea there was to train a separate neural network to mimic the behavior of a thermodynamic solver. You generate a bunch of pseudo-data using a model like UNIFAC, train the surrogate to predict phase compositions from a Gibbs energy curve, and then freeze that surrogate and use it during training of the main model.
Tom: And what's wrong with that? It sounds like a reasonable workaround.
Jane: It has a couple of big problems. First, the surrogate is not thermodynamically consistent. It's just a neural network that has learned to approximate a mapping. It can give you predictions that violate the extremum principle, or even the mass balance. It's not guaranteed to be the global minimum of anything.
Tom: So it's a black box that can give you physically impossible answers.
Jane: Exactly. And the second problem is data leakage. The surrogate is trained on UNIFAC data, but UNIFAC itself was parameterized using experimental data from many of the same molecules that appear in the test set. So the surrogate has implicitly seen the test set before.
Tom: That's a subtle but critical flaw. It's not cheating in the obvious sense, but it's definitely giving the baseline an unfair advantage.
Jane: And despite that advantage, the DISCOMAX method still wins. On the single-system fitting task, the best DISCOMAX variant got an MAE of zero point zero one six, while the best surrogate variant got zero point one zero nine. That's a six-fold improvement.
Tom: And that's just for fitting one system. What about the full generalization task?
Jane: There, DISCOMAX with a Hessian loss got a test MAE of zero point zero six eight, compared to zero point zero eight zero for the surrogate. It's a smaller gap, but still a clear win, and it's achieved without the data leakage problem.
Tom: So the improvements are not just about accuracy. It's about the fundamental approach. DISCOMAX is thermodynamically consistent by construction, it doesn't require any synthetic training data for a surrogate, and it's genuinely end-to-end trainable.
Jane: And it's more robust. The paper shows that the surrogate solver can converge smoothly to entirely wrong solutions when fitting individual systems. The gradients are stable, but they're biased. They don't point toward the true equilibrium.
Tom: That's a scary thought. You think you're training well, but you're actually just getting better at being wrong.
Jane: Exactly. And the DISCOMAX approach avoids that because the forward pass is always the true global minimum on the grid. The gradients are always pointing in a direction that improves the physical consistency of the model.
Tom: I also noticed they introduced a novel Hessian loss to promote convexity at the phase compositions and concavity at the feed. That's a clever auxiliary loss that helps the model learn the right curvature.
Jane: It does, and it's a nice addition. It's not strictly necessary, but it helps, especially for systems with narrow miscibility gaps.
Tom: So, to sum up, this paper improves on the previous work by being more accurate, more physically consistent, and more principled. It's a clear step forward.
Jane: And it opens up a lot of possibilities. Let's think about what this means for the future.
Conclusion: Tom: Well, we've covered a lot of ground on "Differentiable Thermodynamic Phase-Equilibria for Machine Learning." Let's try to wrap our heads around what this really means.
Jane: For me, the biggest takeaway is that we now have a template for embedding rigorous physics into machine learning, even when the target quantity is defined implicitly as the solution of a complex optimization problem.
Tom: And that's not just for chemical engineering. This straight-through estimator idea, combined with a discrete enumeration of feasible states, could be applied to any problem where you need to differentiate through an argmin or argmax.
Jane: Sure, but let's not forget the immediate impact. For chemical engineers, this is a game-changer. It means they can train models directly on liquid-liquid equilibrium data without having to rely on surrogate models that might give physically inconsistent results.
Tom: And it's not just LLE. The paper shows case studies for vapor pressure calculation and even a ternary three-phase equilibrium. The framework is general.
Jane: Right. And the fact that they're releasing the full dataset and code is a huge deal for reproducibility. Anyone can now build on this work.
Tom: I think the biggest impact will be in process design. If you can predict phase behavior accurately and consistently, you can design better distillation columns, extraction units, and crystallization processes.
Jane: And that has real-world implications for making chemical processes more efficient and sustainable. Less energy, less waste, better products.
Tom: There are still limitations, of course. The brute-force grid search scales exponentially with the number of components. For a binary system, it's fast. For a ten-component mixture, it would be computationally prohibitive.
Jane: That's a fair point. But the authors acknowledge that, and they point out that most experimental data is for binary systems anyway. And for deployment, you can always switch to a more traditional flash calculation method.
Tom: So it's a practical tool for training, even if it's not the tool you'd use for every single prediction at scale.
Jane: Exactly. And I think the most exciting future direction is incorporating data on miscible systems. Right now, they only train on systems that exhibit a phase split. If they could also train on single-phase systems, the model could learn to predict whether a phase split occurs at all.
Tom: That would be the holy grail. A model that can tell you both the number of phases and their compositions.
Jane: It would be. And this paper lays the groundwork for exactly that.
Tom: Well, on that note, I think we've given this paper a proper send-off. It's a brilliant piece of work, and we're excited to see where it leads.
Jane: Absolutely. Thanks for joining us, everyone. We'll be back soon with another paper to dissect. Take care.
Tom: See you next time.
Karim K. Ben Hicham, Moreno Ascani, Jan G. Rittig, Alexander Mitsos
RWTH Aachen University · Forschungszentrum Jülich GmbH · JARA-ENERGY
cs.LG
Submitted: 2026-08-17
Updated: 2026-08-18
Comments: 45 pages, 27 figures, 5 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: "We address the challenge of integrating rigorous thermodynamic constraints into machine-learning models whose target quantities are defined implicitly as solutions of an extremum principle.
Key concepts
- DISCOMAX
- An algorithm that solves thermodynamic equilibrium problems by searching a discrete grid of possible compositions. It finds the true minimum Gibbs energy in the forward pass and uses a masked softmax to provide smooth gradients for training neural networks during the backward pass.
- Thermodynamic Consistency
- The principle that machine learning predictions must obey physical laws, such as mass balance and energy minimization. Unlike surrogate models that might provide physically impossible answers, this method ensures predictions are physically valid by construction.
- Straight-through Gradient Estimator
- A technique used to train neural networks when the target function is not differentiable. It allows the model to use a "hard" physical solution for predictions while using a "soft," smooth approximation to calculate the gradients necessary for learning.
Terminology
Summary
Summary
The paper introduces DISCOMAX, a differentiable algorithm for phase-equilibrium calculation that guarantees thermodynamic consistency at both training and inference, only subject to a user-specified discretization. The method combines discrete enumeration of feasible phase states with masked softmax aggregation in the backward pass, with the propagation of the true equilibrium state in the forward pass, using a straight-through gradient estimator to enable physics-consistent end-to-end learning of neural excess Gibbs energy (g E)-models.
The authors state: "We address the challenge of integrating rigorous thermodynamic constraints into machine-learning models whose target quantities are defined implicitly as solutions of an extremum principle. Focusing on binary LLE, we introduce a differentiable equilibrium solver that reformulates the flash calculation as a discrete minimization problem over the ∆g mix surface. Thermodynamic consistency is enforced exactly in the forward pass by evaluating the hard minimizer, while differentiability is maintained in the backward pass via a straight-through gradient estimator that propagates gradients via a Boltzmann-weighted aggregation over discretized compositions."
The paper notes that "the situation changes fundamentally when the quantity of interest is itself defined implicitly as the solution of an equation or even as the minimizer of a constrained optimization problem, especially in the presence of many solutions. In such cases, the learning task naturally assumes the structure of a bilevel optimization problem. The authors highlight that
integrating such solvers directly into gradient-based learning frameworks is challenging, and training thermodynamically consistent models on phase-equilibrium data, such as equilibrium compositions of multiphase systems, requires solving a lower-level optimization problem for each data point at every training epoch and differentiating through its solution."
The authors critique the surrogate-based approach of Hoffmann et al. (2025), noting two critical limitations: "First, the surrogate solver is thermodynamically inconsistent: a learned direct mapping from discretized g E or Gibbs energy of mixing (∆g mix) values to phase compositions does not guarantee satisfaction of any of the following fundamental requirements: (i) global optimality of ∆g mix, (ii) stability conditions, even in the limit of arbitrarily fine discretization, and (iii) a mass-balance if the feed composition z is constrained. Secondly, the surrogate training data are, in our opinion, still biased towards the test set because they are generated from the mod. UNIFAC model, whose parameterization was fitted using many of the molecules that appear in the test set."
The paper draws an analogy to statistical thermodynamics: Phase equilibrium thermodynamics is fundamentally probabilistic: any macroscopic observable O at equilibrium corresponds to an ensemble average over a large number of accessible microscopic states i.
The authors note that "the softmax function... has the same mathematical structure as a Boltzmann-weighted distribution over discrete energy states. Introducing a temperature-like parameter τ (which, in the Boltzmann distribution, is equivalent to the kB T term) controls the sharpness of this distribution and recovers the classical equilibrium solution in the limit τ → 0, while yielding a regular and differentiable mapping for finite τ."
The neural parameterization uses a graph neural network (GNN) encoder: Each molecule pair (mol1 and mol2), is encoded using a graph neural network (GNN), which maps the corresponding molecular graphs to continuous embeddings.
The model predicts ∆g mix (x; θ) = x(1 − x) geE (x; M, θ) + ∆g id (x), where geE is the unrestricted network prediction and ∆g id (x) = x ln(x) + (1 − x) ln(1 − x) is the ideal Gibbs energy of mixing.
The term x(1 − x) guarantees ∆g mix (0) = ∆g mix (1) = 0.
The equilibrium prediction is formulated as: (x′∗, x′′∗) ∈ arg min ∆g mix (x′; θ)ϕ′ + ∆g mix (x′′; θ)ϕ′′,
where x′ and x′′ denote phase composition in phases one and two, respectively, and the molar fractions ϕ′, ϕ′′ of both phases are given by
ϕ′ = (x′′ − z)/(x′′ − x′) and ϕ′′ = (z − x′)/(x′′ − x′).
The DISCOMAX algorithm works as follows: Define an augmented grid by appending the feed composition... Compute the vector of Gibbs energies on the grid... For each pair of grid indices (i, j) with 1 ≤ i ≤ j ≤ N ′, compute
the total Gibbs energy of mixing, where infeasible pairs (those not satisfying mass balance) are assigned +∞. Determine hard minimum indices for inference (not differentiable)... Set x′∗ = x(i∗), x′′∗ = x(j∗).
Then, Compute softmax weights: p(i,j) = exp(−∆g mix,(i,j) /τ) / Σ Σ exp(−∆g mix,(u,v) /τ)
and Compute soft estimates for training (differentiable)
via weighted averages. Finally, Straight-Through Gradient Estimation
is applied: x′∗ ← x′∗ + (x′∗ soft − stopgrad(x′∗ soft))
and similarly for x′′∗.
The authors explain: "Forward pass (fp) (thermodynamically exact, up to discretization). We compute the hard minimizer... This guarantees that the reported equilibrium state satisfies the extremum principle and the mass-balance feasibility constraints encoded in the masking of infeasible pairs. Backward pass (bp) (smooth surrogate gradients). At the same time, we form a differentiable softmax distribution over all candidate pairs... Gradients are then propagated through these soft estimates using a straight-through update."
The loss function includes direct MSE losses L1 = ∥x̂′ − x′∗ ∥22 and L2 = ∥x̂′′ − x′′∗ ∥22, plus two auxiliary losses. The Hessian loss promotes thermodynamically consistent stability behavior at the measured data points, specifically, convexity at the target phase compositions (x′∗, x′′∗) and concavity at the feed composition (z).
The masked Gibbs loss promotes the correct qualitative curvature of the molar Gibbs energy of mixing ∆gmix (T, x) over the composition domain for data points known to exhibit an LLE.
The dataset contains 8,597 unique binary systems with two liquid phases
formed from all possible binary mixtures of 526 unique solvents, resulting in 138,601 candidate systems.
The ground truth equilibrium data is obtained by evaluating the HANNA 2 model on a 401-point grid of compositions.
In the single-system fitting experiments, "the best-performing configurations are DISCOMAX H+G with an MAE of 0.016 ± 0.017 and DISCOMAX H with 0.017 ± 0.018, compared to 0.109 ± 0.064 for Surrogate H and 0.109 ± 0.071 for Surrogate H+G, the strongest surrogate variants. The best surrogate variant has an approximately 6.8× higher MAE relative to DISCOMAX H+G. The authors note:
Even the pure DISCOMAX method without any auxiliary loss performs about 5.7× better than the best surrogate methods and about 13× better than the previously proposed Surrogate G method, in terms of MAE."
The authors observe: "The surrogate solver consistently displays substantially larger deviations from the specified target equilibrium compositions and frequently produces thermodynamically inconsistent equilibrium composition predictions, with violations of the extremum principle, i.e., the predicted configuration is unstable, and not even metastable. Even the satisfaction of simple physical laws, like the conservation of mass, cannot be guaranteed for the surrogate solver."
In the generalization experiments with 10-fold cross-validation, "the DISCOMAX solver augmented with the Hessian-based stability loss (DISCOMAX H) achieves the best test performance, with a mean absolute error of 0.068 ± 0.003, a root mean squared error of 0.090±0.005, and an R2 of 0.970±0.003. The surrogate solver reaches a test MAE of 0.080±0.002, while the plain DISCOMAX solver remains slightly more accurate with a test MAE 0.074±0.003."
The authors discuss: "While the surrogate solver performs unexpectedly poorly when fitted to individual systems, these deficiencies do not fully propagate to the full-dataset training scenario. A likely explanation is that fitting each system in isolation exposes the surrogate solver to its own inaccurate equilibrium estimates, causing the optimization to stabilize at incorrect solutions."
The paper concludes: "Our results demonstrate that rigorous equilibrium thermodynamics can be embedded into modern machine-learning pipelines when the target properties arise from fairly small lower-level optimization problems. The proposed framework establishes a principled foundation for learning thermodynamically grounded models directly from phase-equilibrium data and can, in principle, be extended to other types of phase behavior, including vapor-liquid and solid-liquid equilibria, as well as to more complex chemical systems."
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement and the resulting capabilities of the improved AI system:
Implementation: Replace non-differentiable argmin operations in phase-equilibrium calculations with a straight-through gradient estimator that combines:
-
Hard forward pass: discrete enumeration of feasible phase states + global minimization of Gibbs energy of mixing
-
Soft backward pass: Boltzmann-weighted softmax aggregation over discretized compositions with temperature annealing (τ starting at 0.1, multiplied by 0.98 per epoch)
Resulting capability: The AI system can now train end-to-end on liquid-liquid equilibrium (LLE) data while guaranteeing thermodynamic consistency (global optimality of Gibbs energy, stability conditions, mass balance) up to user-specified discretization precision—something no previous ML approach achieved.
Implementation: Enforce thermodynamic consistency at pure-component limits by multiplying the neural network output by x(1−x) before adding the ideal Gibbs energy term, ensuring Δg mix(0) = Δg mix(1) = 0 exactly. Use smooth ELU activations in the composition-dependent module to guarantee differentiability with respect to composition.
Implementation: Add a loss term that penalizes:
-
Non-convexity of Δg mix at target phase compositions (necessary for stability)
-
Non-concavity at the feed composition (promotes correct miscibility gap location)
This loss uses a margin ϵ to avoid oscillations and is evaluated via automatic differentiation of the Hessian.
Implementation: Eliminate the two-stage approach (training a surrogate solver on UNIFAC data, then using it for gE model training). Instead, integrate the differentiable solver directly into the training loop, avoiding data leakage from test-domain molecules.
Implementation: Use the DISCOMAX solver with Hessian loss, which shows stable gradient norms and robust performance across batch sizes (16–512), unlike the surrogate baseline that diverges at higher learning rates.
-
Predict liquid-liquid equilibrium compositions for binary mixtures with mean absolute error of 0.068 ± 0.003 (vs 0.080 for the best surrogate baseline), achieving R2 = 0.970 on unseen test systems.
-
Fit individual LLE systems exactly with MAE of 0.016 ± 0.017 (vs 0.109 for the best surrogate variant), a 6.8× improvement, while guaranteeing thermodynamic consistency in every prediction.
-
Handle narrow miscibility gaps (width < 0.20) with significantly lower error and variance than surrogate approaches, which struggle in this regime.
-
Scale to full datasets of 8,597 binary LLE systems with ten-fold cross-validation, using a graph neural network (Chemprop D-MPNN) encoder for molecular structure.
-
Extend to other equilibrium types including:
-
Pure-component vapor pressure calculations (demonstrated: 0.4% error on propane)
-
Ternary three-phase liquid-liquid-liquid equilibria (demonstrated on water-dodecane-ionic liquid system)
-
Potentially solid-liquid equilibria and multicomponent systems
-
Provide thermodynamically consistent predictions at inference time—every predicted phase composition corresponds to the global minimum of the Gibbs energy of mixing, satisfying mass balance and stability conditions, unlike surrogate-based approaches that can produce thermodynamically impossible results.
-
Train without auxiliary losses if needed—the plain DISCOMAX variant still outperforms the best surrogate baseline (MAE 0.074 vs 0.080), making the method robust to loss-function choices.
Abstract
Accurate prediction of phase equilibria remains a central challenge in chemical engineering. Physics-consistent machine learning methods that incorporate thermodynamic structure into neural networks have recently shown strong performance for activity-coefficient modeling. However, extending such approaches to equilibrium data arising from an extremum principle, such as liquid-liquid equilibria, remains difficult. Here we present DISCOMAX, a differentiable algorithm for phase-equilibrium calculation that guarantees thermodynamic consistency at both training and inference, only subject to a user-specified discretization. The method combines discrete enumeration of feasible phase states with masked softmax aggregation in the backward pass, with the propagation of the true equilibrium state in the forward pass, using a straight-through gradient estimator to enable physics-consistent end-to-end learning of neural-models. We show that this approach bears analogy to statistical thermodynamics, and we evaluate it on binary liquid-liquid equilibrium data where it outperforms existing surrogate-based methods, while offering a general framework for learning from different kinds of equilibrium data.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks