Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Scaling Machine Learning Interatomic Potentials with Mixtures of Experts".
Jane: The paper was written by Yuzhi Liu, Duo Zhang, Anyang Peng, Weinan E, Linfeng Zhang et al. from AI for Science Institute and DP Technology and Peking University and Institute of Applied Physics and Computational Mathematics.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the computational chemistry world, and it’s called “Scaling Machine Learning Interatomic Potentials with Mixtures of Experts.”
Jane: And Tom, I have to say, the title alone tells you exactly what the stakes are. We’ve seen this “mixture of experts” idea completely transform large language models, and now these researchers are asking whether the same trick can make atomistic simulations smarter and faster.
Tom: Right, and for our listeners who might not live and breathe this stuff, let’s break down what an interatomic potential even is. It’s basically a mathematical model that tells you how much energy a group of atoms has based on their positions.
Jane: Exactly. And if you can predict that energy accurately, you can simulate how molecules fold, how materials deform, how catalysts work — all without running expensive quantum mechanics calculations every single step.
Tom: The problem is, the more accurate you want that potential to be, the bigger and more expensive the neural network behind it gets. And that’s where this paper comes in with a clever solution.
Jane: So instead of one giant dense network that has to process every atom with all of its parameters, you build a bunch of smaller “expert” networks. A router decides which experts are relevant for each atom, and you only activate a few of them at a time.
Tom: It’s like having a team of specialists instead of one generalist. You don’t need the cardiologist to look at your broken ankle, right? You just route that case to the bone doctor.
Jane: And that’s the core promise here. You get the capacity of a huge model, but you only pay the computational cost of a small one. The paper is testing whether that promise actually holds up when you’re dealing with continuous potential energy surfaces.
Tom: Which is trickier than it sounds, because in language models, you can switch experts between words without anyone noticing. But in physics, if the energy function jumps or becomes jagged when you switch experts, you break the laws of energy conservation.
Jane: That’s the real tension this paper tackles. They’re not just bolting on a known technique — they’re adapting it to the unique demands of physical simulation.
Tom: And the authors here are a serious crew. Linfeng Zhang and Han Wang are basically royalty in this field — they built the Deep Potential framework that’s been used to simulate everything from battery materials to proteins.
Jane: So they have the credibility to push this forward. The question is whether their specific design choices actually beat the standard approach, and that’s what we’re going to dig into next.
Tom: Stay with us, because the results are genuinely surprising — especially when it comes to how they route the experts.
Summary: Jane: So Tom, we’ve set the stage. Now let’s talk about what this paper actually found when they put their mixture-of-experts model to the test.
Tom: And the headline is that it works — really well. They built their system on top of a model called DPA3, which is already one of the best interatomic potentials out there, and they showed that adding this expert architecture gives you a serious accuracy boost.
Jane: Right, and the numbers are pretty striking. On the OMol25 benchmark, which covers eighty-three chemical elements and all sorts of messy real-world chemistry, their best configuration cut the energy error by about twenty-six percent compared to the baseline DPA3 model.
Tom: And that’s not just a small tweak. They’re also beating a dense model that has six times the parameters. So you’re getting better accuracy with less compute — that’s the dream scenario.
Jane: But what I find really interesting is that they didn’t just throw a bunch of experts at the problem and hope for the best. They systematically tested different design choices, and some of their findings are counterintuitive.
Tom: Like what?
Jane: Well, they found that if you have a pool of experts but no “shared” experts — meaning every expert is only activated for specific elements — the model stops improving once you get past a certain number of activated experts. It plateaus.
Tom: But when they added shared experts — a few that are always on, capturing knowledge common across all elements — the scaling just kept going. The model got better and better as they increased the expert pool.
Jane: It makes sense when you think about it. Every atom, regardless of what element it is, follows the same laws of physics. There’s a common core of behavior — how atoms repel each other at short range, how bonds stretch — and that should be handled by a shared set of experts.
Tom: Then the specialized experts can focus on what makes each element unique, like how a heavy metal behaves differently from a light gas.
Jane: Exactly. And they also found that how you route the experts matters enormously. They tested two strategies: one where the router looks at each individual atom and its element type, and another where the router looks at the whole configuration and picks one set of experts for everyone.
Tom: The global routing approach sounds simpler, but it actually failed catastrophically in some cases. The model just wouldn’t train — the errors blew up and never converged.
Jane: Which is a huge red flag, because that’s the approach another major model called UMA used. This paper is essentially showing that element-wise routing is not just better — it’s necessary for stable training.
Tom: And when they looked at why the element-wise routing works so well, they found something beautiful. The router’s learned representations of different elements organize themselves along the periodic table.
Jane: So the model is essentially rediscovering chemistry. It clusters alkali metals together, transition metals together, and the patterns it learns mirror the trends chemists have known for a century.
Tom: That’s the kind of result that makes you sit up and pay attention. The model isn’t just memorizing — it’s finding structure that aligns with fundamental science.
Jane: And that structure is why it generalizes so well across different datasets. They tested it on molecular data, solid-state materials, and catalytic surface reactions, and it beat the baseline everywhere.
Tom: So the summary is: sparse experts plus shared experts plus element-wise routing equals a better, more interpretable model. And we haven’t even talked about the trickiest part yet.
Jane: Which is?
Tom: How do you keep the energy surface smooth when you’re switching experts between atoms? That’s the engineering puzzle we’re going to dig into next.
Improvements: Jane: Welcome back. Tom just teased the big engineering challenge, and I think this is where the paper really earns its keep.
Tom: So here’s the problem. In a language model, if you switch from one expert to another between two tokens, nobody cares. The output is just text — it doesn’t need to be smooth or continuous.
Jane: But in an interatomic potential, the energy surface has to be smooth. If you move an atom a tiny bit and suddenly the model switches which experts are active, the energy could jump. That jump means the forces become infinite or undefined, and your simulation just explodes.
Tom: And the paper’s solution to this is clever. They make the routing depend only on the atomic number — the element type — not on the atom’s position. So as an atom moves, its expert assignment stays fixed.
Jane: That guarantees smoothness, because the gating weights don’t change as coordinates change. The only thing changing is the input features that go into the fixed set of experts.
Tom: But that creates a different problem. If routing depends only on element type, then every carbon atom in the system uses the same experts, regardless of whether it’s in a diamond, a protein, or a carbon nanotube.
Jane: And that’s where the shared experts come back in. The shared experts capture the universal physics, and the element-specific experts capture the chemical identity. Together, they can represent the full range of environments without needing the router to be position-dependent.
Tom: Right. And the paper also compares two ways of combining the experts. One is called MoE, where each expert is its own little neural network with its own activation function, and you mix their outputs.
Jane: The other is called MoLE, where you linearly combine the expert weights first and then apply one activation function at the end. It’s like averaging the “personalities” of the experts before the network does its thing.
Tom: And the surprising result is that the simpler-looking MoLE actually underperforms the more complex MoE when shared experts are present. The nonlinearity inside each expert matters — it lets each expert specialize more sharply.
Jane: It’s like the difference between having a team where each member gives their full opinion and you synthesize them, versus having everyone whisper a few words and then you try to guess what they meant. The first approach preserves more information.
Tom: And the improvements aren’t marginal. On the OMol25 dataset, their best MoE model with shared experts achieved a normalized force error of about zero point seven zero, meaning it’s thirty percent better than the baseline. That’s a massive jump.
Jane: They also showed that the MoE model beats a dense model with the same activated parameter count. So it’s not just about having more parameters — it’s about how you structure them.
Tom: And here’s the kicker for the engineering-minded folks: the MoE model with sixty-four total experts but only six activated per atom uses the same compute as a model with six times fewer parameters. You get the accuracy of a huge model at the cost of a small one.
Jane: That’s the kind of efficiency gain that could make these simulations accessible to labs that don’t have access to massive GPU clusters.
Tom: And when they visualized what the experts were doing, they found that the router had essentially learned the periodic table on its own. Elements in the same group — like alkali metals — got similar routing patterns.
Jane: That’s not just a nice visualization. It means the model is allocating its capacity in a way that aligns with chemical intuition, which makes it more trustworthy and easier to debug.
Tom: So the improvements here are threefold: smooth routing via element-wise gating, better capacity via shared experts, and sharper specialization via nonlinear expert networks.
Jane: And the result is a model that’s not only more accurate but also more interpretable. That’s a rare combination in deep learning.
Tom: We’ve got one more segment to wrap this up, and I want to talk about what this means for the future of atomistic simulation.
Conclusion: Jane: Alright, we’re in the home stretch. Let’s pull together what we’ve learned from “Scaling Machine Learning Interatomic Potentials with Mixtures of Experts.”
Tom: The big picture is that this paper gives the community a clear recipe for scaling up atomistic models without hitting the wall of diminishing returns that dense networks run into.
Jane: And the recipe has three ingredients: sparse activation so you only pay for what you use, shared experts to capture universal physics, and element-wise routing to keep the energy surface smooth while letting each element have its own specialists.
Tom: They validated this on three major benchmarks — molecules, materials, and catalysts — and beat both the baseline and a much larger dense model. That’s a strong signal that this isn’t a fluke.
Jane: And the interpretability result — the router spontaneously organizing elements along periodic trends — suggests the model is learning something fundamental about chemistry, not just memorizing training data.
Tom: Now, I want to bring in Lu and Meng, because they’ve been listening and I know they have thoughts.
Lu: I’m excited about what this means for foundation models in chemistry. We’re seeing the same pattern that played out in language models — sparse experts let you scale to trillions of parameters without blowing up the compute budget. This paper is the first serious step toward that for interatomic potentials.
Meng: And from an engineering standpoint, the fact that they didn’t need to change the training infrastructure much is huge. You can take an existing DPA3 model, swap the dense layers for MoE layers, and get a better model with the same compute. That lowers the barrier to adoption.
Jane: That’s a good point, Meng. The paper is honest about one limitation though — they haven’t yet built the fully distributed training system that would let experts live on different GPUs and communicate efficiently.
Tom: Right, but the architecture is inherently parallelizable. Once you have expert parallelism, you can scale to much larger models and datasets without hitting communication bottlenecks.
Lu: And that’s where I think the real impact will come. Imagine training a single model on every known material and molecule, with experts that specialize not just by element but by chemical environment. That could give us a truly universal interatomic potential.
Meng: The other thing I appreciate is that they’re not overselling it. They show where it works well and where the gains are smaller — like on the solid-state OMat24 dataset, the improvement was more modest. That honesty helps people know where to apply this.
Jane: So what’s the takeaway for our listeners? If you’re working on molecular dynamics, drug discovery, materials design, or catalysis, this paper gives you a concrete way to get more accuracy from your models without buying more GPUs.
Tom: And for the field as a whole, it points toward a future where atomistic simulations are powered by large, sparse, interpretable models — the same way large language models have transformed text processing.
Jane: We’re going to say goodbye to this paper now, but I suspect we’ll be seeing its ideas show up in a lot of follow-up work over the next year.
Tom: Thanks for joining us, everyone. Next up, we’ve got a paper on equivariant diffusion models for protein design, so stay tuned.
Jane: See you then.
Yuzhi Liu, Duo Zhang, Anyang Peng, Weinan E, Linfeng Zhang, Han Wang
AI for Science Institute · DP Technology · Peking University · Institute of Applied Physics and Computational Mathematics
physics.chem-ph, cs.LG, physics.comp-ph
Submitted: 2026-08-17
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: "Here we systematically develop Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPs and analyze the effects of routing strategies and expert designs." The work
Key concepts
- Interatomic Potential
- A mathematical model that predicts the energy of a group of atoms based on their positions. Accurate potentials allow simulations to model molecular folding or material deformation without running expensive quantum mechanics calculations at every step.
- Mixture of Experts (MoE)
- An architecture where a large model is replaced by several smaller 'expert' networks. A router decides which experts are relevant for each input, allowing the system to use only a few specialized networks at a time, saving computational cost.
- Element-wise Routing
- A method where the routing mechanism depends only on the atomic number or element type of an atom, not its specific position. This ensures that as an atom moves, its expert assignment remains fixed, which helps maintain a smooth energy surface during simulation.
Terminology
Summary
Summary
The paper Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
by Yuzhi Liu, Duo Zhang, Anyang Peng, Weinan E, Linfeng Zhang, and Han Wang, presents a systematic investigation of Mixture-of-Experts (MoE) architectures for Machine Learning Interatomic Potentials (MLIPs). The authors state: Here we systematically develop Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPs and analyze the effects of routing strategies and expert designs.
The work addresses two fundamental challenges in applying MoE to MLIPs: first, standard MoE designs are often incompatible with the equivariant representations employed by MLIP architectures based on equivariant graph neural networks (GNN), such as MACE, SevenNet, eSEN, and AlphaNet
; second, "unlike the discrete token representations in language models, where sparse expert activation is natural, MLIPs model continuous potential energy surfaces. Abrupt expert switching in this setting can introduce numerical instabilities or non-physical discontinuities, potentially violating the law of energy conservation."
The authors propose an MoE-integrated DPA3 MLIP architecture, noting that DPA3 employs exclusively invariant node and edge features, which permits valid nonlinear operations on expert outputs and thus enables the seamless incorporation of standard MoE mechanisms within its framework.
They introduce element-wise expert gating, yielding an element-dependent MoE that models potential-energy contributions through species-specific expert routing
and incorporate a shared-expert mechanism, ensuring that a subset of experts is always activated to capture knowledge common across all element types.
The MoE transformation is defined mathematically: for atom i, the output is the sum of gated routed experts plus shared experts, where gating scores are computed via softmax over a routing matrix applied to a latent representation of the atom's chemical identity. The hidden representation is defined as ui = MLP(one hot(Zi))
, where Zi is the atomic number. The authors distinguish between element-wise routing (MoE-E), where each atom has its own gating scores, and global routing (MoE-G), where all atoms share an identical set of gating scores and weights
based on a mean-pooled representation.
The MoLE formulation is described as mathematically equivalent to a single dense layer with effective parameters
where the features are first linearly transformed by expert-specific weights, after which the expert contributions are aggregated via gating weights, and a nonlinear activation function is applied at the end.
Key experimental findings on the OMol25 dataset include:
-
Sparse activation and shared experts:
In the absence of shared experts, the normalized energy (force) MAEs for K = 2, 4, 6, and 8 are 1.082 (0.965), 0.965 (0.919), 0.923 (0.902), and 0.926 (0.900), respectively.
The authors note thatperformance saturates between K = 6 and K = 8, indicating diminishing returns.
However,introducing shared experts consistently improves model performance,
andperformance is optimized when the number of shared experts approaches approximately half of the activated experts.
With half shared experts,the normalized energy (force) MAEs for K = 2, 4, 6, and 8 are reduced to 0.939 (0.913), 0.771 (0.779), 0.740 (0.747), and 0.673 (0.715), respectively,
showing thatthe accuracy improvement in this regime does not plateau at K = 8, but instead continues to scale favorably.
-
MoE-E vs MoLE-E:
Both models show comparable performance without shared experts,
with discrepancies within 0.02 and 0.03 for energy and force MAE. However, with shared experts,MoE-E demonstrates a substantially stronger response,
achievinga larger reduction in normalized energy MAE, ranging from 0.18 to 0.23, and a decrease in normalized force MAE of approximately 0.13 to 0.14,
while MoLE-E shows reductions of approximately 0.11 in energy and 0.06-0.09 in force. -
Parameter efficiency: The MoE-E model with four activated experts
consistently outperforms the widened baseline
(a dense model with 4× parameters), achievingadditional gains of 0.052 and 0.048 in energy and force precision
with 64 experts. -
Element-wise vs Global routing:
The MoE-G architecture suffers from catastrophic training failure: its energy and force MAEs exhibit severe numerical instability and fail to converge.
For stable variants,MoE-E consistently outperforms both MoLE variants across all expert scales,
andMoLE-G systematically underperforms its element-wise counterpart MoLE-E, with MoLE-E achieving approximately 5%–10% lower MAEs.
On multi-dataset benchmarks (OMol25, OMat24, OC20M), the MoE-E model (with K=6 activated experts, 3 shared, N=64 total) consistently outperforms MoLE-E across all evaluated benchmarks,
with reductions in normalized energy MAE of 0.10, 0.02, and 0.03
and force MAE of 0.10, 0.03, and 0.04
respectively. Furthermore, MoE-E even surpasses a dense baseline with 6× the parameter count, yielding additional error reductions of up to 0.07 in Energy and 0.05 in Force.
The authors also analyze expert weighting distribution patterns using PCA on the OMat24 dataset, finding that the first two principal components accounting for 93.57% of the total explained variance.
They observe that lanthanides and actinides cluster on the left side of the PCA manifold, while transition metals predominantly occupy the central region,
and that elements within the same group arrange themselves diagonally from the top-left to the bottom-right of the PCA space,
with lighter elements located in the top-left region and heavier elements positioned toward the bottom-right.
The authors conclude that the MoE-E router effectively encodes elemental characteristics in a manner consistent with fundamental physicochemical principles.
The paper concludes: sparse activation, shared-expert mechanisms, nonlinear expert specialization, and element-wise expert routing collectively constitute the key ingredients enabling stable and scalable performance gains.
The authors acknowledge a limitation: the current implementation does not yet exploit a fully distributed training and inference framework specifically optimized for sparse expert parallelism,
and state that future work will therefore focus on developing large-scale distributed MoE training and inference pipelines.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Replace dense feed-forward layers in graph neural network-based interatomic potential models (e.g., DPA3) with a sparse MoE architecture featuring:
-
Element-wise routing based on atomic number (via one-hot encoding → MLP)
-
Shared experts (approximately half of activated experts)
-
Nonlinear activation within each expert before mixing
Resulting capability: The model achieves 10-20% lower energy/force prediction errors compared to dense baselines with equivalent parameter counts, while maintaining computational efficiency through sparse activation.
Abstract
Machine Learning Interatomic Potentials (MLIPs) enable accurate large-scale atomistic simulations, yet improving their expressive capacity efficiently remains challenging. Here we systematically develop Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPs and analyze the effects of routing strategies and expert designs. We show that sparse activation combined with shared experts yields substantial performance gains, and that nonlinear MoE formulations outperform MoLE when shared experts are present, underscoring the importance of nonlinear expert specialization. Furthermore, element-wise routing consistently surpasses configuration-level routing, while global MoE routing often leads to numerical instability. The resulting element-wise MoE model achieves state-of-the-art accuracy across the OMol25, OMat24, and OC20M benchmarks. Analysis of routing patterns reveals chemically interpretable expert specialization aligned with periodic-table trends, indicating that the model effectively captures element-specific chemical characteristics for precise interatomic modeling.
Sources
- Directional Message Passing for Molecular Graphs
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- A Graph Neural Network for the Era of Large Atomistic Models
- Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- UMA: A Family of Universal Models for Atoms
- From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property Prediction
- What Does BERT Look At? An Analysis of BERT's Attention
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- ST-MoE: Designing Stable and Transferable Sparse Expert Models
- Efficient Large Scale Language Modeling with Mixtures of Experts
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
- Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models
Related papers
- Transferable Generative Models Bridge Femtosecond to Nanosecond Time-Step Molecular Dynamics
- Accelerated "on-the-fly" coupled-cluster path-integral molecular dynamics: Impact of nuclear quantum effects on an asymmetric proton
- Variational Polaron Theory for Ground States of Strongly Coupled Light-Matter and Electron-Phonon Systems
- Pushing the accuracy of on-top functionals with agent-driven supervised learning
- Localized intrinsic bond orbitals decode correlated charge migration dynamics
- Smite: A quasiclassical trajectory (QCT) program for bimolecular collisions and unimolecular dynamics on ab-initio, and machine-learned potential energy surfaces