When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling
Yong Yi Bay, Kathleen A. Yearick
cs.LG, cs.AI, cs.CL, stat.ML
Submitted: 2026-06-27
Comments: 24 pages, 10 figures, 3 tables. Code and data: https://github.com/bay-yearick-lab/sampling-ceilings
Code: https://github.com/bay-yearick-lab/sampling-ceilings
License: http://creativecommons.org/licenses/by/4.0/
The gist: People overthink; language models over-sample, and the extra effort can talk both into a worse answer.
Terminology
Abstract
People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sampling it many times (test-time scaling), and the more they draw, the more often a correct answer turns up somewhere, so coverage, the fraction of problems with at least one correct try, climbs and appears to be progress. But a deployed system must return one answer, and choosing it, not knowing which try is right, is selection; selection is capped, and past a point extra samples only make the model surer of a confident mistake, even as every draw adds cost. The gap between climbing coverage and stalled selection, the identifiability gap, is the answer a model can produce but not pick. So the real question is not whether to sample but how far, and the answer is: not far. For picking an answer, the vote has already settled within a few dozen draws, the modal ceiling; for scoring a benchmark, sooner still, the correlation ceiling. Beyond that, extra draws cost compute and add nothing, and can even make the answer worse. This paper turns the cutoff into a single number, the effective number of samples, that any sampling run already reveals. The bottleneck is recognizing a right answer, not generating one.
Sources
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- OpenAI o1 System Card
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- s1: Simple test-time scaling
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Evaluating Large Language Models Trained on Code
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
- Great Models Think Alike and this Undermines AI Oversight
- Model Capability Dominates: Inference-Time Optimization Lessons from AIMO 3
- How Do Large Language Monkeys Get Their Power (Laws)?
- Efficient Prediction of Pass@k Scaling in Large Language Models
- A Simple Model of Inference Scaling Laws
- Training Verifiers to Solve Math Word Problems
- No 3D Matrices: A Unified Tensor-Product View of Matrix-Free Cartesian PDE Solvers
- Machine Learning vs Deep Learning: The Generalization Problem
- Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
- On the Effect of Sampling Diversity in Scaling LLM Inference
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks