When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems
cs.AI, cs.LG
Submitted: 2026-09-13
Updated: 2026-09-13
Comments: 35 pages, 3 figures, 9 tables
License: http://creativecommons.org/licenses/by/4.0/
The gist: AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction.
Terminology
Abstract
AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a higher score under a larger budget does not by itself show where additional resources are best spent. This critical integrative review asks when a reported scaling result supports a resource-allocation decision. It compares evidence across pretraining, test-time computation, retrieval, and agent evaluation, distinguishing the performance of a tested procedure from the best performance achievable under a resource limit. The synthesis shows that three mismatches recur across this evidence: success counted before an answer is chosen, information a deployed system will not have, and costs left out of the comparison. A capability surface expresses performance as a function of budgets, mechanisms, and available information. Worked analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can alter an allocation conclusion. A resource envelope provides a structured record of the task, development and run-time resources, information access, and procedure behind a reported score. Its application to a published comparison illustrates which conclusions the evidence supports and which deployment questions remain unresolved. The resulting framework specifies the comparisons needed to choose among feasible systems and motivates experiments on the transfer of allocation rules across tasks and operating conditions. It does not propose a universal scaling law or infer general intelligence from benchmark gains.
Sources
- Explaining Neural Scaling Laws
- Improving language models by retrieving from trillions of tokens
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Accelerating Large Language Model Decoding with Speculative Sampling
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Optimizing Model Selection for Compound AI Systems
- Evaluating Large Language Models Trained on Code
- On the Measure of Intelligence
- Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning
- Training Verifiers to Solve Math Word Problems
- Rational Metareasoning for Large Language Models
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- Reusing Pre-Training Data at Test Time is a Compute Multiplier
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Datasheets for Datasets
- Adaptive Computation Time for Recurrent Neural Networks
- REALM: Retrieval-Augmented Language Model Pre-Training
- Deep Learning Scaling is Predictable, Empirically
- Distilling the Knowledge in a Neural Network
- Training Compute-Optimal Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection