ADAPT: Lightweight, Long-Range Machine Learning Force Fields Without Graphs
cs.LG, cond-mat.mtrl-sci, math.OC
Submitted: 2025-09-28
Updated: 2026-09-15
Comments: 14 total pages of main content, 4 of references, 3 in Appendix
Journal ref: Phys. Rev. Research 8, 033266 (2026)
DOI: 10.1103/3xf3-ts5f
Project page: https://evandramko.github.io/files/attention.pdf
License: http://creativecommons.org/licenses/by/4.0/
The gist: Point defects play a central role in driving the properties of materials.
Terminology
Abstract
Point defects play a central role in driving the properties of materials. First-principles methods are widely used to compute defect energetics and structures, including at scale for high-throughput defect databases. However, these methods are computationally expensive, making machine-learning force fields (MLFFs) an attractive alternative for accelerating structural relaxations. Most existing MLFFs are based on graph neural networks (GNNs), which can suffer from oversmoothing and poor representation of long-range interactions. Both of these issues are especially of concern when modeling point defects. To address these challenges, we introduce the Accelerated Deep Atomic Potential Transformer (ADAPT), an MLFF that replaces graph representations with a direct coordinates-in-space formulation and explicitly considers all pairwise atomic interactions. Atoms are treated as tokens, with a Transformer encoder modeling their interactions. Applied to a dataset of silicon point defects, ADAPT achieves a roughly 33 percent reduction in both force and energy prediction errors relative to a state-of-the-art GNN-based model, while requiring only a fraction of the computational cost.
Sources
- Geometric Deep Learning: Grids, Groups, Graphs, Geodesics, and Gauges
- MatterSim: A Deep Learning Atomistic Model Across Elements, Temperatures and Pressures
- So3krates: Equivariant attention for interactions on arbitrary length-scales in molecular systems
- Tutorial on amortized optimization
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Neural Regression For Scale-Varying Targets
- On the space-time expressivity of ResNets
- Approximation properties of Residual Neural Networks for Kolmogorov PDEs
- Advanced Physics-Informed Neural Network with Residuals for Solving Complex Integral Equations
- A foundation model for atomistic materials chemistry
- Orb: A Fast, Scalable Neural Network Potential
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree Representations
- The dark side of the forces: assessing non-conservative force models for atomistic machine learning
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Flamingo: a Visual Language Model for Few-Shot Learning
- Training language models to follow instructions with human feedback
- Understanding Why ViT Trains Badly on Small Datasets: An Intuitive Perspective
- When Do You Need Billions of Words of Pretraining Data?
- Graph Neural Networks for Relational Inductive Bias in Vision-based Deep Reinforcement Learning of Robot Control
- A Survey on Oversmoothing in Graph Neural Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks