Hamilton-Zero: A Neural Tensor-Network Foundation Model for Ground States of Arbitrary Quadratic Qubit Hamiltonians

arXiv:2608.11911 · quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI · Submitted 2026-08-20 · Read on arXiv

Timothy Heightman, Elena Orlova, Philip Mantrov, Aleksei Ustimenko

ICFO – Institute of Photonic Sciences · Simulacra Research Inc.

quant-ph, cond-mat.dis-nn, cond-mat.str-el, cs.AI

Submitted: 2026-08-20

Updated: 2026-08-24

Comments: 22 pages main text

Code: https://github.com/simulacra-research/HamiltonZero

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 100/100

The gist: Hamilton-Zero is a foundation model for computing ground states of arbitrary quadratic qubit Hamiltonians, trained on a dataset of hundreds of thousands of Hamiltonian systems.

Terminology

Summary

Hamilton-Zero is a foundation model for computing ground states of arbitrary quadratic qubit Hamiltonians, trained on a dataset of hundreds of thousands of Hamiltonian systems. The model has approximately 0.5 billion variational parameters and is trained using techniques from large language models and deep reinforcement learning. The key innovation is formulating spin-1/2 quantum ground-state learning as manifold variational optimisation over centrally odd scalar functions on SU(2)N, replacing explicit Hilbert-space vector amplitudes with manifold functions on which the Hamiltonian acts through Lie derivatives, evaluated by custom automatic differentiation primitives. The paper proves that the resulting variational principle on this manifold preserves the spin-1/2 sector's ground-state upper bound using the Peter-Weyl theorem.

The model is pretrained on a dataset of 5000 different interaction topologies spanning a century of quantum many-body literature, with perturbations and exact symmetry transformations expanding this to hundreds of thousands of distinct training Hamiltonians. The architecture consists of a Hamiltonian featurizer, a trunk of attention and FFN layers, a leaf builder where spin configurations enter, and a merge-tree readout with a learnable routing policy trained with deep reinforcement learning to infer entanglement structure.

Results show that a single shared optimisation trajectory improves the energy of a heterogeneous distribution of Hamiltonians, with the V-score falling approximately as C-0.53 and the ED-relative gap as C-0.61. On held-out generalisation datasets, zero-shot inference achieves median signed gaps of 4.02%, 4.21%, and 12.63% on three evaluation sets, with fine-tuning reducing errors by two to three orders of magnitude. The model can be taken far outside its training system-size distribution, evaluating systems up to 8100 qubits. It can witness phase transitions from a single checkpoint of its weights, becomes approximately equivariant with respect to gauge groups, and the router effectively learns entanglement structure that generalises to unseen topologies and sizes.

The paper also details significant engineering contributions, including a novel SU(2) replica-exchange Langevin sampler, sharded natural-gradient optimisation with an extension of the KFAC optimiser, and custom automatic differentiation primitives that scale to 8000+ qubits. The compute scaling shows approximately 5× better scaling than today's LLMs. The model's weights are released open-source, and the paper argues this changes the economics of quantum ground-state computation, with classical amortized computation now serving as the baseline for claims of useful quantum advantage.

Improvements for AI systems

Improvements to AI systems:

  1. Manifold-based variational inference for high-dimensional physical systems – Replace fixed-vector state representations with learnable manifold functions (e.g., on SU(2) N) where the objective acts via Lie derivatives. This enables continuous, differentiable optimization over exponentially large state spaces without explicit Hilbert-space storage.

  2. Automatic differentiation primitives for Lie-group actions – Implement custom AD kernels that compute Lie derivatives and group-equivariant gradients directly on manifold coordinates, scaling to 8000+ parameters while preserving exactness. This allows any neural network to optimize over compact Lie groups efficiently.

  3. Reinforcement-learned routing for hierarchical readouts – Use a policy network to dynamically select which sub-branches of a tree-structured output to activate based on input structure (e.g., entanglement topology). This generalizes to unseen graph sizes and connectivity patterns, improving compositional generalization.

  4. Pretraining on physics-derived data with symmetry augmentation – Train on a diverse set of interaction topologies (from many-body literature) with exact symmetry transformations (gauge, permutation, time-reversal) as data augmentation. This yields approximate equivariance and zero-shot transfer to new Hamiltonians, reducing fine-tuning cost by 2–3 orders of magnitude.

  5. Sharded natural-gradient optimization with KFAC extension – Extend KFAC to non-Euclidean parameter manifolds (e.g., SU(2) tensor products) with sharded computation across devices. This provides stable, curvature-aware updates for 0.5B-parameter models, achieving 5× better compute scaling than standard LLM training.

  6. Replica-exchange Langevin sampling on Lie groups – Implement a parallel tempering sampler that exchanges replicas across different inverse temperatures on SU(2) N, enabling robust exploration of rugged loss landscapes and avoiding local minima in variational quantum problems.


What the improved AI system can do:

  • Solve arbitrary quantum many-body ground states (spin-1/2, arbitrary topology) with a single forward pass, achieving median energy errors <5% zero-shot and <0.1% after minimal fine-tuning, for systems up to 8100 qubits.

  • Witness quantum phase transitions from a single checkpoint of weights, without retraining, by analyzing the router’s learned entanglement structure across parameter sweeps.

  • Generalize to unseen interaction graphs and sizes – e.g., from 2D lattices to random graphs or 3D topologies – without architectural changes, due to the RL-routed merge-tree and symmetry-augmented pretraining.

  • Optimize any cost function defined on a compact Lie group (not just quantum Hamiltonians) – e.g., in robotics (SO(3) orientation control), protein folding (torsion angles), or combinatorial optimization – using the same manifold variational framework and AD primitives.

  • Serve as a reusable foundation model for physics – released open-source – where users can fine-tune on their specific Hamiltonian family with only a few hundred examples, replacing expensive exact diagonalization or Monte Carlo baselines.

  • Scale to 8000+ dimensional continuous optimization with natural-gradient stability, enabling training of larger models on limited hardware via sharded KFAC and replica-exchange sampling.

Sources

Related papers