PMCTS: Principled Parallelized Inference Time Scaling with Particle Monte Carlo Tree Search
cs.LG
Submitted: 2026-05-09
Updated: 2026-09-07
Code: https://github.com/ax-ml/jax
License: http://creativecommons.org/licenses/by/4.0/
The gist: Monte Carlo Tree Search (MCTS) is a widely used approach for policy improvement and action selection in Reinforcement Learning.
Terminology
Abstract
Monte Carlo Tree Search (MCTS) is a widely used approach for policy improvement and action selection in Reinforcement Learning. Due to its sequential and deterministic nature, principled runtime-scaling of MCTS with parallel compute remains a major challenge. We introduce Particle MCTS (PMCTS), a principled parallel MCTS algorithm suited for neural network evaluations and designed for GPU-acceleration with batch-parallelization. We establish policy improvement guarentees for modern MCTS algorithms and show that PMCTS maintains them. Empirically, PMCTS scales well with parallel compute and consistently outperforms or compares well to the popular heuristic-based baselines across a range of MCTS and RL evaluation domains, including the board games chess and Go and popular discrete action and continuous control benchmarks.
Sources
- MuZero with Self-competition for Rate Control in VP9 Video Compression
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
- On Effective Parallelization of Monte Carlo Tree Search
- TransZero: Parallel Tree Expansion in MuZero using Transformer Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks