More Bang for the Buck: Improving the Inference of Large Language Models at a Fixed Budget using Reset and Discard (ReD)
cs.LG, cond-mat.dis-nn, cs.AI, stat.ML
Submitted: 2026-01-29
Updated: 2026-09-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials.
Terminology
Abstract
The performance of large language models (LLMs) on verifiable tasks is usually measured by pass@k, the probability of answering a question correctly at least once in k trials. At a fixed budget, a more suitable metric is coverage@cost, the average number of unique questions answered as a function of the total number of attempts. We connect the two metrics and show that the empirically-observed power-law behavior in pass@k leads to a sublinear growth of the coverage@cost (diminishing returns). To solve this problem, we propose Reset-and-Discard (ReD), a query method of LLMs that increases coverage@cost for a given budget, regardless of the pass@k form. Moreover, given a pass@k, we can quantitatively predict the savings in the total number of attempts using ReD. If pass@k is not available for the model, ReD can infer its power-law exponent. Experiments on three LLMs across coding (HumanEval), math (GSM8K), and reasoning (MMLU-Pro) benchmarks demonstrate that ReD substantially reduces the required attempts, tokens, and USD cost to reach a desired coverage, while also offering an efficient way to measure inference power-laws. ReD's advantage is maintained for imperfect verifiers and outperforms the tested allocation baselines.
Sources
- Evaluating Large Language Models Trained on Code
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
- Baldur: Whole-Proof Generation and Repair with Large Language Models
- Best-of-N Jailbreaking
- Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
- How Do Large Language Monkeys Get Their Power (Laws)?
- CodeMonkeys: Scaling Test-Time Compute for Software Engineering
- RoboMonkey: Scaling Test-Time Sampling and Verification for Vision-Language-Action Models
- Efficient Prediction of Pass@k Scaling in Large Language Models
- From Decoding to Meta-Generation: Inference-time Algorithms for Large Language Models
- Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling
- SGDR: Stochastic Gradient Descent with Warm Restarts
- gpt-oss-120b & gpt-oss-20b Model Card
- Reinforced Self-Training (ReST) for Language Modeling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks