Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time Scaling
cs.LG, cs.DC, cs.PF
Submitted: 2026-09-16
Updated: 2026-09-19
Comments: 8 pages, 4 figures, 9 tables. Experiments evaluate Phi-3-mini and Qwen2.5-1.5B on GSM8K and SciQ using NVIDIA A100 and V100 GPUs. Studies LLM test-time scaling, candidate-generation scheduling, latency, throughput, GPU-hours, and GPU-device energy
License: http://creativecommons.org/licenses/by/4.0/
The gist: Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses.
Terminology
Abstract
Test-time scaling can improve large language model reasoning by generating and combining multiple candidate responses. In sampling-based methods, the inference budget is often described by the number of generated candidates, N. However, N tells us how many candidates are generated, not how they are executed. The same candidate budget can be produced in one batched generation call or split across several sequential calls with smaller batch sizes. We first study the effect of increasing N on reasoning accuracy using Phi-3-mini and Qwen2.5-1.5B on 500 GSM8K prompts. As expected, increasing N from 1 to 8 improves accuracy by 8.4 percentage points for Phi-3-mini and 18.4 points for Qwen2.5-1.5B. However, accuracy alone does not show the systems cost of using a larger candidate budget. We therefore fix N = 8 and compare four generation schedules: 1x8, 2x4, 4x2, and 8x1, where axb denotes a generation calls with b candidates per call. We measure latency, throughput, GPU-hours, and gross GPU-device energy while keeping the total candidate count fixed. On A100 GPUs, eight serial calls use 4.64-4.86x as much gross GPU-device energy and have 5.77-6.12x the P95 latency of one batched call with eight candidates. The same pattern appears across three independently scheduled A100 nodes per model and in short-output SciQ/V100 experiments. These results show that candidate count alone is not enough to describe the systems cost of multi-candidate test-time scaling. When candidates are independent and memory allows it, fewer generation calls with larger batch sizes are more efficient. Evaluations should therefore report not only candidate count and accuracy, but also generation schedule and GPU-level systems metrics.
Sources
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
- Universal Self-Consistency for Large Language Model Generation
- BEST-Route: Adaptive LLM Routing with Test-Time Optimal Compute
- Reliable Chain-of-Thought via Prefix Consistency
- Interpretable Adaptive Sampling for LLM Test-Time Scaling
- PerLLM: Personalized Inference Scheduling with Edge-Cloud Collaboration for Diverse LLM Services
- Online Scheduling for LLM Inference with KV Cache Constraints
- How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference
- Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
- Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
- Energy Considerations of Large Language Model Inference and Efficiency Optimizations
- Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
- Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2.5 Technical Report
- Training Verifiers to Solve Math Word Problems
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks