IsingFormer: Augmenting Parallel Tempering With Learned Proposals
cs.LG, cond-mat.stat-mech, cs.AI, physics.comp-ph
Submitted: 2025-09-27
Updated: 2026-09-11
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative models have been extensively used to accelerate MCMC mixing for sampling and optimization, but their effective integration with standard MCMC remains an open question.
Terminology
Abstract
Generative models have been extensively used to accelerate MCMC mixing for sampling and optimization, but their effective integration with standard MCMC remains an open question. Here, we introduce a global proposal move in which finite-temperature configurations from an external generator are used as proposals within Parallel Tempering (PT). We examine a specific generator, IsingFormer, a Transformer trained on long-run MCMC configurations intended to approximate equilibrium distributions, and call the resulting framework Transformer-Augmented Parallel Tempering (TAPT). The IsingFormer exhibits two useful capabilities: interpolation to untrained β values and conditional completion under clamped settings absent from training. On 3D spin-glass instances, TAPT reaches substantially lower residual energies than standard PT in fewer Monte Carlo sweeps. On integer factorization, we show how a structured problem encoding can amortize the training cost by training IsingFormer once and reusing the same model across target products that were not imposed during training. Finally, in a scaling study that uses long-run MCMC configurations as proposals, with proposal generation excluded from the timing, TAPT reduces the fitted time-to-solution exponent by approximately 33% relative to PT over the tested problem sizes.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks