Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-25
Code: https://github.com/Tencent-Hunyuan/Hyra-results
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time.
Terminology
Abstract
Large language models (LLMs) can improve solutions to verifiable scientific and algorithmic problems by spending additional computation at test time. Recent systems achieve strong results with increasingly elaborate evolutionary search harnesses or by updating model parameters during test-time training. We ask how much of this machinery is necessary. We introduce Hill Sampling, a simple procedure that repeatedly samples candidate program edits from a frozen LLM, retains the best program found so far, and conditions all subsequent samples on that program. We evaluate the method on circle packing, sums/differences of sets, and Erdos' minimum-overlap problem using three open-weight models. Hill Sampling sets a new state of the art on circle packing among published methods, improves over the AlphaEvolve reference on Erdos' minimum-overlap problem, and achieves strong results on sums and differences of finite sets. The circle-packing and Erdos results require only hours of wall-clock time on eight NVIDIA H100 GPUs. To our knowledge, we also conduct, the largest study, by parameter count, of evolution strategies (ES) applied directly to LLM weights at test time. Surprisingly, learning the weights is worse than setting the ES learning rate to zero: at zero learning rate, the method is still searching in weight space through fixed random perturbations. Those perturbations can help exploration, but randomness from token sampling is stronger still, and repeated sampling remains substantially weaker than Hill Sampling. These results suggest a simple test-time compute allocation strategy: repeatedly sample edits to the best verified solution found so far, before introducing additional complexity such as adding archives, diversity mechanisms, evolutionary scaffolds, or test-time parameter learning.
Sources
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization
- Sums and differences of sets (improvement over AlphaEvolve)
- Automated Discovery Has No Universally Superior Harness
- Settling the Optimal Exponent Relating Sumsets and Difference Sets
- NGRPO: Negative-enhanced Group Relative Policy Optimization
- AlphaEvolve: A coding agent for scientific and algorithmic discovery
- Evolution Strategies as a Scalable Alternative to Reinforcement Learning
- Structured Scaling of AI Discovery Across Diverse Scientific Domains
- Sums and differences of sets: a further improvement over AlphaEvolve
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks