Scaling Machine Learning Interatomic Potentials with Mixtures of Experts
summary
The gist
"Here we systematically develop Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPs and analyze the effects of routing strategies and expert designs." The work
In short
The episode discusses a paper scaling machine learning interatomic potentials using Mixtures of Experts. The hosts explain how this method uses smaller 'expert' networks routed by an element-wise mechanism to achieve better accuracy and efficiency than dense models. Key findings include improved performance, structural organization mirroring the periodic table, and a recipe for scaling atomistic simulations.
Key concepts
- Interatomic Potential
- A mathematical model that predicts the energy of a group of atoms based on their positions. Accurate potentials allow simulations to model molecular folding or material deformation without running expensive quantum mechanics calculations at every step.
- Mixture of Experts (MoE)
- An architecture where a large model is replaced by several smaller 'expert' networks. A router decides which experts are relevant for each input, allowing the system to use only a few specialized networks at a time, saving computational cost.
- Element-wise Routing
- A method where the routing mechanism depends only on the atomic number or element type of an atom, not its specific position. This ensures that as an atom moves, its expert assignment remains fixed, which helps maintain a smooth energy surface during simulation.
Terminology used across episodes
This episode discusses
- Scaling Machine Learning Interatomic Potentials with Mixtures of Experts · Paper Radio
- Directional Message Passing for Molecular Graphs
- Scaling Laws for Neural Language Models
- Training Compute-Optimal Large Language Models
- A Graph Neural Network for the Era of Large Atomistic Models
- Learning Smooth and Expressive Interatomic Potentials for Physical Property Prediction
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- UMA: A Family of Universal Models for Atoms
- From Molecules to Materials: Pre-training Large Generalizable Models for Atomic Property Prediction
- What Does BERT Look At? An Analysis of BERT's Attention
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- ST-MoE: Designing Stable and Transferable Sparse Expert Models
- Efficient Large Scale Language Modeling with Mixtures of Experts
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- The Open Molecules 2025 (OMol25) Dataset, Evaluations, and Models
- Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models
The paper
Scaling Machine Learning Interatomic Potentials with Mixtures of Experts · Read on arXiv
Yuzhi Liu, Duo Zhang, Anyang Peng, Weinan E, Linfeng Zhang, Han Wang
AI for Science Institute · DP Technology · Peking University · Institute of Applied Physics and Computational Mathematics
Machine Learning Interatomic Potentials (MLIPs) enable accurate large-scale atomistic simulations, yet improving their expressive capacity efficiently remains challenging. Here we systematically develop Mixture-of-Experts (MoE) and Mixture-of-Linear-Experts (MoLE) architectures for MLIPs and analyze the effects of routing strategies and expert designs. We show that sparse activation combined with shared experts yields substantial performance gains, and that nonlinear MoE formulations outperform MoLE when shared experts are present, underscoring the importance of nonlinear expert specialization. Furthermore, element-wise routing consistently surpasses configuration-level routing, while global MoE routing often leads to numerical instability. The resulting element-wise MoE model achieves state-of-the-art accuracy across the OMol25, OMat24, and OC20M benchmarks. Analysis of routing patterns reveals chemically interpretable expert specialization aligned with periodic-table trends, indicating that the model effectively captures element-specific chemical characteristics for precise interatomic modeling.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Scaling Machine Learning Interatomic Potentials with Mixtures of Experts".
Jane: The paper was written by Yuzhi Liu, Duo Zhang, Anyang Peng, Weinan E, Linfeng Zhang et al. from AI for Science Institute and DP Technology and Peking University and Institute of Applied Physics and Computational Mathematics.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the computational chemistry world, and it’s called “Scaling Machine Learning Interatomic Potentials with Mixtures of Experts.”
Jane: And Tom, I have to say, the title alone tells you exactly what the stakes are. We’ve seen this “mixture of experts” idea completely transform large language models, and now these researchers are asking whether the same trick can make atomistic simulations smarter and faster.
Tom: Right, and for our listeners who might not live and breathe this stuff, let’s break down what an interatomic potential even is. It’s basically a mathematical model that tells you how much energy a group of atoms has based on their positions.
Jane: Exactly. And if you can predict that energy accurately, you can simulate how molecules fold, how materials deform, how catalysts work — all without running expensive quantum mechanics calculations every single step.
Tom: The problem is, the more accurate you want that potential to be, the bigger and more expensive the neural network behind it gets. And that’s where this paper comes in with a clever solution.
Jane: So instead of one giant dense network that has to process every atom with all of its parameters, you build a bunch of smaller “expert” networks. A router decides which experts are relevant for each atom, and you only activate a few of them at a time.
Tom: It’s like having a team of specialists instead of one generalist. You don’t need the cardiologist to look at your broken ankle, right? You just route that case to the bone doctor.
Jane: And that’s the core promise here. You get the capacity of a huge model, but you only pay the computational cost of a small one. The paper is testing whether that promise actually holds up when you’re dealing with continuous potential energy surfaces.
Tom: Which is trickier than it sounds, because in language models, you can switch experts between words without anyone noticing. But in physics, if the energy function jumps or becomes jagged when you switch experts, you break the laws of energy conservation.
Jane: That’s the real tension this paper tackles. They’re not just bolting on a known technique — they’re adapting it to the unique demands of physical simulation.
Tom: And the authors here are a serious crew. Linfeng Zhang and Han Wang are basically royalty in this field — they built the Deep Potential framework that’s been used to simulate everything from battery materials to proteins.
Jane: So they have the credibility to push this forward. The question is whether their specific design choices actually beat the standard approach, and that’s what we’re going to dig into next.
Tom: Stay with us, because the results are genuinely surprising — especially when it comes to how they route the experts.
Summary: Jane: So Tom, we’ve set the stage. Now let’s talk about what this paper actually found when they put their mixture-of-experts model to the test.
Tom: And the headline is that it works — really well. They built their system on top of a model called DPA3, which is already one of the best interatomic potentials out there, and they showed that adding this expert architecture gives you a serious accuracy boost.
Jane: Right, and the numbers are pretty striking. On the OMol25 benchmark, which covers eighty-three chemical elements and all sorts of messy real-world chemistry, their best configuration cut the energy error by about twenty-six percent compared to the baseline DPA3 model.
Tom: And that’s not just a small tweak. They’re also beating a dense model that has six times the parameters. So you’re getting better accuracy with less compute — that’s the dream scenario.
Jane: But what I find really interesting is that they didn’t just throw a bunch of experts at the problem and hope for the best. They systematically tested different design choices, and some of their findings are counterintuitive.
Tom: Like what?
Jane: Well, they found that if you have a pool of experts but no “shared” experts — meaning every expert is only activated for specific elements — the model stops improving once you get past a certain number of activated experts. It plateaus.
Tom: But when they added shared experts — a few that are always on, capturing knowledge common across all elements — the scaling just kept going. The model got better and better as they increased the expert pool.
Jane: It makes sense when you think about it. Every atom, regardless of what element it is, follows the same laws of physics. There’s a common core of behavior — how atoms repel each other at short range, how bonds stretch — and that should be handled by a shared set of experts.
Tom: Then the specialized experts can focus on what makes each element unique, like how a heavy metal behaves differently from a light gas.
Jane: Exactly. And they also found that how you route the experts matters enormously. They tested two strategies: one where the router looks at each individual atom and its element type, and another where the router looks at the whole configuration and picks one set of experts for everyone.
Tom: The global routing approach sounds simpler, but it actually failed catastrophically in some cases. The model just wouldn’t train — the errors blew up and never converged.
Jane: Which is a huge red flag, because that’s the approach another major model called UMA used. This paper is essentially showing that element-wise routing is not just better — it’s necessary for stable training.
Tom: And when they looked at why the element-wise routing works so well, they found something beautiful. The router’s learned representations of different elements organize themselves along the periodic table.
Jane: So the model is essentially rediscovering chemistry. It clusters alkali metals together, transition metals together, and the patterns it learns mirror the trends chemists have known for a century.
Tom: That’s the kind of result that makes you sit up and pay attention. The model isn’t just memorizing — it’s finding structure that aligns with fundamental science.
Jane: And that structure is why it generalizes so well across different datasets. They tested it on molecular data, solid-state materials, and catalytic surface reactions, and it beat the baseline everywhere.
Tom: So the summary is: sparse experts plus shared experts plus element-wise routing equals a better, more interpretable model. And we haven’t even talked about the trickiest part yet.
Jane: Which is?
Tom: How do you keep the energy surface smooth when you’re switching experts between atoms? That’s the engineering puzzle we’re going to dig into next.
Improvements: Jane: Welcome back. Tom just teased the big engineering challenge, and I think this is where the paper really earns its keep.
Tom: So here’s the problem. In a language model, if you switch from one expert to another between two tokens, nobody cares. The output is just text — it doesn’t need to be smooth or continuous.
Jane: But in an interatomic potential, the energy surface has to be smooth. If you move an atom a tiny bit and suddenly the model switches which experts are active, the energy could jump. That jump means the forces become infinite or undefined, and your simulation just explodes.
Tom: And the paper’s solution to this is clever. They make the routing depend only on the atomic number — the element type — not on the atom’s position. So as an atom moves, its expert assignment stays fixed.
Jane: That guarantees smoothness, because the gating weights don’t change as coordinates change. The only thing changing is the input features that go into the fixed set of experts.
Tom: But that creates a different problem. If routing depends only on element type, then every carbon atom in the system uses the same experts, regardless of whether it’s in a diamond, a protein, or a carbon nanotube.
Jane: And that’s where the shared experts come back in. The shared experts capture the universal physics, and the element-specific experts capture the chemical identity. Together, they can represent the full range of environments without needing the router to be position-dependent.
Tom: Right. And the paper also compares two ways of combining the experts. One is called MoE, where each expert is its own little neural network with its own activation function, and you mix their outputs.
Jane: The other is called MoLE, where you linearly combine the expert weights first and then apply one activation function at the end. It’s like averaging the “personalities” of the experts before the network does its thing.
Tom: And the surprising result is that the simpler-looking MoLE actually underperforms the more complex MoE when shared experts are present. The nonlinearity inside each expert matters — it lets each expert specialize more sharply.
Jane: It’s like the difference between having a team where each member gives their full opinion and you synthesize them, versus having everyone whisper a few words and then you try to guess what they meant. The first approach preserves more information.
Tom: And the improvements aren’t marginal. On the OMol25 dataset, their best MoE model with shared experts achieved a normalized force error of about zero point seven zero, meaning it’s thirty percent better than the baseline. That’s a massive jump.
Jane: They also showed that the MoE model beats a dense model with the same activated parameter count. So it’s not just about having more parameters — it’s about how you structure them.
Tom: And here’s the kicker for the engineering-minded folks: the MoE model with sixty-four total experts but only six activated per atom uses the same compute as a model with six times fewer parameters. You get the accuracy of a huge model at the cost of a small one.
Jane: That’s the kind of efficiency gain that could make these simulations accessible to labs that don’t have access to massive GPU clusters.
Tom: And when they visualized what the experts were doing, they found that the router had essentially learned the periodic table on its own. Elements in the same group — like alkali metals — got similar routing patterns.
Jane: That’s not just a nice visualization. It means the model is allocating its capacity in a way that aligns with chemical intuition, which makes it more trustworthy and easier to debug.
Tom: So the improvements here are threefold: smooth routing via element-wise gating, better capacity via shared experts, and sharper specialization via nonlinear expert networks.
Jane: And the result is a model that’s not only more accurate but also more interpretable. That’s a rare combination in deep learning.
Tom: We’ve got one more segment to wrap this up, and I want to talk about what this means for the future of atomistic simulation.
Conclusion: Jane: Alright, we’re in the home stretch. Let’s pull together what we’ve learned from “Scaling Machine Learning Interatomic Potentials with Mixtures of Experts.”
Tom: The big picture is that this paper gives the community a clear recipe for scaling up atomistic models without hitting the wall of diminishing returns that dense networks run into.
Jane: And the recipe has three ingredients: sparse activation so you only pay for what you use, shared experts to capture universal physics, and element-wise routing to keep the energy surface smooth while letting each element have its own specialists.
Tom: They validated this on three major benchmarks — molecules, materials, and catalysts — and beat both the baseline and a much larger dense model. That’s a strong signal that this isn’t a fluke.
Jane: And the interpretability result — the router spontaneously organizing elements along periodic trends — suggests the model is learning something fundamental about chemistry, not just memorizing training data.
Tom: Now, I want to bring in Lu and Meng, because they’ve been listening and I know they have thoughts.
Lu: I’m excited about what this means for foundation models in chemistry. We’re seeing the same pattern that played out in language models — sparse experts let you scale to trillions of parameters without blowing up the compute budget. This paper is the first serious step toward that for interatomic potentials.
Meng: And from an engineering standpoint, the fact that they didn’t need to change the training infrastructure much is huge. You can take an existing DPA3 model, swap the dense layers for MoE layers, and get a better model with the same compute. That lowers the barrier to adoption.
Jane: That’s a good point, Meng. The paper is honest about one limitation though — they haven’t yet built the fully distributed training system that would let experts live on different GPUs and communicate efficiently.
Tom: Right, but the architecture is inherently parallelizable. Once you have expert parallelism, you can scale to much larger models and datasets without hitting communication bottlenecks.
Lu: And that’s where I think the real impact will come. Imagine training a single model on every known material and molecule, with experts that specialize not just by element but by chemical environment. That could give us a truly universal interatomic potential.
Meng: The other thing I appreciate is that they’re not overselling it. They show where it works well and where the gains are smaller — like on the solid-state OMat24 dataset, the improvement was more modest. That honesty helps people know where to apply this.
Jane: So what’s the takeaway for our listeners? If you’re working on molecular dynamics, drug discovery, materials design, or catalysis, this paper gives you a concrete way to get more accuracy from your models without buying more GPUs.
Tom: And for the field as a whole, it points toward a future where atomistic simulations are powered by large, sparse, interpretable models — the same way large language models have transformed text processing.
Jane: We’re going to say goodbye to this paper now, but I suspect we’ll be seeing its ideas show up in a lot of follow-up work over the next year.
Tom: Thanks for joining us, everyone. Next up, we’ve got a paper on equivariant diffusion models for protein design, so stay tuned.
Jane: See you then.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language