Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians
Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko
ICFO – Institute of Photonic Sciences · Simulacra Research Inc.
quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI
Submitted: 2026-08-20
Updated: 2026-08-24
Comments: 22 pages main text
Code: https://github.com/simulacra-research/HamiltonZero
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: Hamilton-Zero is a foundation model for computing ground states of arbitrary quadratic qubit Hamiltonians, trained on a dataset of hundreds of thousands of Hamiltonian systems.
Terminology
Summary
Hamilton-Zero is a foundation model for computing ground states of arbitrary quadratic qubit Hamiltonians, trained on a dataset of hundreds of thousands of Hamiltonian systems. The model has approximately 0.5 billion variational parameters and is trained using techniques from large language models and deep reinforcement learning. The key innovation is formulating spin-1/2 quantum ground-state learning as manifold variational optimisation over centrally odd scalar functions on SU(2)N, replacing explicit Hilbert-space vector amplitudes with manifold functions on which the Hamiltonian acts through Lie derivatives, evaluated by custom automatic differentiation primitives. The paper proves that the resulting variational principle on this manifold preserves the spin-1/2 sector's ground-state upper bound using the Peter-Weyl theorem.
The model is pretrained on a dataset of 5000 different interaction topologies spanning a century of quantum many-body literature, with perturbations and exact symmetry transformations expanding this to hundreds of thousands of distinct training Hamiltonians. The architecture consists of a Hamiltonian featurizer, a trunk of attention and FFN layers, a leaf builder where spin configurations enter, and a merge-tree readout with a learnable routing policy trained with deep reinforcement learning to infer entanglement structure.
Results show that a single shared optimisation trajectory improves the energy of a heterogeneous distribution of Hamiltonians, with the V-score falling approximately as C-0.53 and the ED-relative gap as C-0.61. On held-out generalisation datasets, zero-shot inference achieves median signed gaps of 4.02%, 4.21%, and 12.63% on three evaluation sets, with fine-tuning reducing errors by two to three orders of magnitude. The model can be taken far outside its training system-size distribution, evaluating systems up to 8100 qubits. It can witness phase transitions from a single checkpoint of its weights, becomes approximately equivariant with respect to gauge groups, and the router effectively learns entanglement structure that generalises to unseen topologies and sizes.
The paper also details significant engineering contributions, including a novel SU(2) replica-exchange Langevin sampler, sharded natural-gradient optimisation with an extension of the KFAC optimiser, and custom automatic differentiation primitives that scale to 8000+ qubits. The compute scaling shows approximately 5× better scaling than today's LLMs. The model's weights are released open-source, and the paper argues this changes the economics of quantum ground-state computation, with classical amortized computation now serving as the baseline for claims of useful quantum advantage.
Improvements for AI systems
Improvements to AI systems:
-
Manifold-based variational inference for high-dimensional physical systems – Replace fixed-vector state representations with learnable manifold functions (e.g., on SU(2) N) where the objective acts via Lie derivatives. This enables continuous, differentiable optimization over exponentially large state spaces without explicit Hilbert-space storage.
-
Automatic differentiation primitives for Lie-group actions – Implement custom AD kernels that compute Lie derivatives and group-equivariant gradients directly on manifold coordinates, scaling to 8000+ parameters while preserving exactness. This allows any neural network to optimize over compact Lie groups efficiently.
-
Reinforcement-learned routing for hierarchical readouts – Use a policy network to dynamically select which sub-branches of a tree-structured output to activate based on input structure (e.g., entanglement topology). This generalizes to unseen graph sizes and connectivity patterns, improving compositional generalization.
-
Pretraining on physics-derived data with symmetry augmentation – Train on a diverse set of interaction topologies (from many-body literature) with exact symmetry transformations (gauge, permutation, time-reversal) as data augmentation. This yields approximate equivariance and zero-shot transfer to new Hamiltonians, reducing fine-tuning cost by 2–3 orders of magnitude.
-
Sharded natural-gradient optimization with KFAC extension – Extend KFAC to non-Euclidean parameter manifolds (e.g., SU(2) tensor products) with sharded computation across devices. This provides stable, curvature-aware updates for 0.5B-parameter models, achieving 5× better compute scaling than standard LLM training.
-
Replica-exchange Langevin sampling on Lie groups – Implement a parallel tempering sampler that exchanges replicas across different inverse temperatures on SU(2) N, enabling robust exploration of rugged loss landscapes and avoiding local minima in variational quantum problems.
What the improved AI system can do:
-
Solve arbitrary quantum many-body ground states (spin-1/2, arbitrary topology) with a single forward pass, achieving median energy errors <5% zero-shot and <0.1% after minimal fine-tuning, for systems up to 8100 qubits.
-
Witness quantum phase transitions from a single checkpoint of weights, without retraining, by analyzing the router’s learned entanglement structure across parameter sweeps.
-
Generalize to unseen interaction graphs and sizes – e.g., from 2D lattices to random graphs or 3D topologies – without architectural changes, due to the RL-routed merge-tree and symmetry-augmented pretraining.
-
Optimize any cost function defined on a compact Lie group (not just quantum Hamiltonians) – e.g., in robotics (SO(3) orientation control), protein folding (torsion angles), or combinatorial optimization – using the same manifold variational framework and AD primitives.
-
Serve as a reusable foundation model for physics – released open-source – where users can fine-tune on their specific Hamiltonian family with only a few hundred examples, replacing expensive exact diagonalization or Monte Carlo baselines.
-
Scale to 8000+ dimensional continuous optimization with natural-gradient stability, enabling training of larger models on limited hardware via sharded KFAC and replica-exchange sampling.
Sources
- Solving the Quantum Many-Body Problem with Artificial Neural Networks
- Approaching the Thermodynamic Limit with Neural-Network Quantum States
- Attention-Based Foundation Model for Quantum States
- Transformer Quantum State: A Multi-Purpose Model for Quantum Many-Body Problems
- Fine-tuning Neural Network Quantum States
- Foundation Neural-Networks Quantum States as a Unified Ansatz for Multiple Hamiltonians
- Quantum Spin Glass in the Two-Dimensional Disordered Heisenberg Model via Foundation Neural-Network Quantum States
- Ab-Initio Solution of the Many-Electron Schr\"odinger Equation with Deep Neural Networks
- Deep neural network solution of the electronic Schr\"odinger equation
- Gold-standard solutions to the Schr\"odinger equation using deep learning: How much physics do we need?
- Large Electron Model: A Universal Ground State Predictor
- QERNEL: a Scalable Large Electron Model
- Solving the electronic Schr\"odinger equation for multiple nuclear geometries with weight-sharing deep neural networks
- Towards a Foundation Model for Neural Network Wavefunctions
- Generalizing Neural Wave Functions
- An ab initio foundation model of wavefunctions that accurately describes chemical bond breaking
- Transformer Wave Function for Quantum Long-Range models
- Neural Scaling Laws Surpass Chemical Accuracy for the Many-Electron Schr\"odinger Equation
- Scaling Laws for Neural-Network Quantum States
- Quantum Machine Learning in Multi-Qubit Phase-Space Part I: Foundations
Related papers
- Reconquering Bell sampling on qudits: stabilizer learning and testing, quantum pseudorandomness bounds, and more
- Encrypted clones can leak: Classification of informative subsets in Quantum Encrypted Cloning
- Polynomial-time classical and quantum simulation of quantum impurity models
- Theory of quantum-enhanced interferometry with general Markovian light sources
- A convergent hierarchy of spectral gap certificates for qubit Hamiltonians
- Universal Bound and Phase Transition in Many-Body Fermionic Non-Gaussianity