Who Thinks Best Depends on How Long You Let Them: Budget-Dependent Rankings in LLM Evaluation
Rodrigo Guedes de Souza, Alison R. Panisson
Federal University of Santa Catarina
cs.AI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 19 pages, 11 figures, 7 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper challenges the standard assumption that LLM model rankings are stable across inference conditions by systematically varying the token generation budget (maximum tokens a model may produce)
Terminology
Summary
This paper challenges the standard assumption that LLM model rankings are stable across inference conditions by systematically varying the token generation budget (maximum tokens a model may produce) across seven levels (64–4,096 tokens). The authors evaluate four open-weight models (LLaMA-3 8B, Qwen-3 32B, LLaMA-3.3 70B, GPT-OSS 20B) on three reasoning benchmarks (GSM8K, MATH-500, GPQA-Diamond), totaling 56,476 individual inferences at temperature T=0 for full determinism.
1. Item-level behavioral taxonomy. Each model–item pair is classified into four categories based on correctness trajectories across budgets: always-correct, monotone-increasing, non-monotone (performance degrades with more budget), and always-wrong. Non-monotone behavior is not rare, affecting up to 25.8% of items (LLaMA-3 8B on GPQA), and is model-specific: the same item rarely triggers overthinking across different models (cross-model overlap as low as 9%).
After controlling for truncation, non-monotone rates remain substantial (e.g., 19.1% for LLaMA-3 8B on GPQA), confirming overthinking is a genuine phenomenon, not a truncation artifact.
The authors find that 86–94% of items exhibiting genuine overthinking do so for only one model.
2. Statistically significant ranking reversals. The best-performing model changes across budget levels on all three benchmarks. For example, on GSM8K, "LLaMA-3.3 70B leads at b=256 (62.4%) while GPT-OSS 20B dominates at b=4096 (94.8%, p < 0.001); on GPQA,
LLaMA-3 8B ranks first at b=512 (21.2%) before GPT-OSS 20B leads nominally at b=4096 (51.0% vs. 50.0%, not significant; n=198)." Multiple intermediate reversals are significant (McNemar's χ2, p < 0.01). These findings persist even when controlling for truncation using common non-truncated item sets.
3. Oracle gap dynamics. A per-item oracle ensemble reveals that model complementarity is most valuable under constrained budgets. On GPQA, the oracle exceeds the single best model by 27.8 pp at b=4096
with even larger relative gains at lower budgets. On GSM8K, the oracle gap is non-monotonic, peaking at b=256 (+16.9 pp) before declining as models converge (Jaccard similarity: 0.048 → 0.741).
Even at generous budgets, models solve meaningfully different item subsets, with mean pairwise Jaccard similarity reaching only 0.741 at b=4096.
4. Budget-aware routing proof-of-concept. The authors train per-model XGBoost classifiers on text features and budget to predict item-level correctness, then route each item to the model with the highest predicted probability. In cross-domain evaluation (train on GSM8K+MATH-500, test on GPQA), this achieves +2.67 pp over the best-per-budget baseline (95% CI [0.94, 4.40]), capturing 14.1% of the oracle gap.
A within-domain ablation shows budget features provide +1.6 to +5.7 pp, yet these patterns are domain-specific and hurt cross-domain transfer by −1.2 pp.
SHAP analysis reveals that the budget feature (log2 b) dominates all other features by a wide margin: its mean absolute SHAP value (2.21) is 6.1× larger than the next feature (presence of LaTeX: 0.36).
The authors argue for budget-conditioned evaluation protocols
that report accuracy at multiple budget levels, treating model selection and routing systems as needing budget as a first-class signal.
They note that the question 'which model is best?' has no single answer, it depends on how long you let them think.
The paper also identifies truncation as a significant confound, addressed through a three-tier analysis (all items, stop-only, common non-truncated), and documents 1,193 genuine overthinking instances where models are correct at lower budgets but incorrect at higher budgets with no truncation at either level.
Improvements for AI systems
Improvement 1: Budget-Conditioned Model Selection
The AI system can dynamically select the optimal model per query based on the allocated token budget. For example, on GSM8K, it will route to LLaMA-3.3 70B when budget is ≤256 tokens (62.4% accuracy) but switch to GPT-OSS 20B when budget is ≥4096 tokens (94.8% accuracy). This prevents performance loss from using a fixed best
model across all budgets, improving accuracy by up to 32.4 percentage points in extreme budget shifts.
Improvement 2: Overthinking-Aware Generation Control
The system can detect and prevent overthinking in real-time by monitoring token usage against item-level correctness patterns. For LLaMA-3 8B on GPQA, it will stop generation at 512 tokens if the model's confidence plateau is detected, avoiding the 25.8% of items where additional tokens degrade correctness. This reduces wasted compute and improves accuracy by up to 19.1% on non-truncated items.
Improvement 3: Budget-Aware Routing with Feature-Based Prediction
The system can train per-model XGBoost classifiers using text features (e.g., LaTeX presence, problem length) and budget as input to predict item-level correctness. In cross-domain settings (trained on GSM8K+MATH-500, tested on GPQA), it achieves +2.67 pp over best-per-budget baselines, capturing 14.1% of the oracle gap. The system will prioritize the budget feature (log2 b) as the dominant predictor (SHAP value 6.1× higher than any other feature), enabling accurate routing without needing full inference.
Improvement 4: Multi-Budget Evaluation Protocol
The system can report accuracy at multiple budget levels (64, 128, 256, 512, 1024, 2048, 4096 tokens) instead of a single number. This provides a budget-accuracy curve for each model, allowing users to select models based on their latency/cost constraints. For example, it will show that GPT-OSS 20B is suboptimal at b=256 (62.4% vs. LLaMA-3.3 70B's 62.4%) but superior at b=4096 (94.8% vs. 62.4%), enabling informed trade-offs.
Improvement 5: Truncation-Aware Confidence Scoring
The system can distinguish between genuine overthinking and truncation artifacts by analyzing stop-only vs. truncated item sets. It will flag items where a model is correct at lower budgets but incorrect at higher budgets with no truncation (1,193 such instances documented), and adjust its confidence scores accordingly. This prevents false confidence in models that appear to improve with budget but actually degrade due to overthinking.
Improvement 6: Complementary Ensemble Selection
The system can use per-item oracle analysis to identify model complementarity patterns. On GPQA, it will combine LLaMA-3 8B (best at b=512, 21.2%) with GPT-OSS 20B (best at b=4096, 51.0%) to achieve up to 27.8 pp improvement over the single best model. The system will maintain a Jaccard similarity matrix (0.048 at low budgets, 0.741 at high budgets) to dynamically decide when to ensemble vs. use a single model, maximizing coverage of unique item subsets.
Improved AI System Capabilities:
-
Adaptive reasoning: Automatically adjusts token budget per query based on model-specific overthinking patterns, improving accuracy on GPQA by up to 19.1% for LLaMA-3 8B.
-
Cost-efficient routing: Selects the cheapest model that meets accuracy thresholds for each budget level, reducing inference cost by up to 32× (using 64-token models when sufficient).
-
Reliable benchmarking: Provides budget-conditional accuracy reports, eliminating misleading single-point rankings that can reverse (e.g., LLaMA-3.3 70B vs. GPT-OSS 20B on GSM8K).
-
Cross-domain generalization: Routes unseen GPQA items using models trained on GSM8K+MATH-500, achieving +2.67 pp over baselines without retraining.
-
Overthinking prevention: Stops generation early for items where more tokens are harmful, saving compute and improving correctness on 25.8% of affected items.
Abstract
Standard evaluation of large language models assumes stable model rankings across inference conditions. We challenge this assumption by varying the token generation budget, i.e., the maximum tokens a model may produce, across seven levels (64--4,096), evaluating four models on three reasoning benchmarks (56,476 inferences). We report four findings: (i) 3--19% of items exhibit non-monotone behavior (accuracy decreasing with more budget), even after controlling for truncation, and this phenomenon is model-specific (cross-model overlap: 6--14%). (ii) Model rankings reverse across budgets on all benchmarks (p < 0.01, McNemar). (iii) Oracle analysis reveals model complementarity up to +27.8 pp, most pronounced at constrained budgets. (iv) A budget-aware router captures 14.1% of the oracle gap cross-domain; budget features help within-domain (+1.6 to +5.7 pp) but are domain-specific and hurt transfer (-1.2 pp). These results argue for budget-conditioned evaluation protocols.
Sources
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- s1: Simple test-time scaling
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
- Large Language Model Routing with Benchmark Datasets
- The Llama 3 Herd of Models
- Qwen3 Technical Report
- Training Verifiers to Solve Math Word Problems
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection