Small-Scale Experiments: Are We There Yet?
Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi
Meta · New York University
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 29 pages, 17 figures
Code: https://github.com/facebookresearch/lingua
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: Small-Scale Experiments: Are We There Yet? Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi Summary This paper addresses the long-standing promise of scaling laws to enable cost-effective
Terminology
Summary
Small-Scale Experiments: Are We There Yet?
Nicholas Lourie, Kyunghyun Cho, Karen Ullrich, Sanae Lotfi
Summary
This paper addresses the long-standing promise of scaling laws to enable cost-effective machine learning experiments by conducting research at small scales and transferring findings to larger scales. The authors argue that this promise has remained unfulfilled because scaling laws are notoriously elusive at small scales,
leading researchers to conclude that sizable models cannot be avoided. However, the paper's central claim is that this is a misconception: the confounding factor is hyperparameters.
The paper demonstrates that scaling laws do emerge at very small scales—as little as 4M parameters
—but only on the fully tuned frontier.
The key finding is that small models are highly sensitive, but hyperparameter sensitivity fades with scale.
This small-scale sensitivity makes scaling laws easy to miss because reaching the tuned frontier requires an extensive search far beyond what most ever run.
The authors show through ablations that well-tuned hyperparameters matter more than any other ingredient
in estimating a scaling law.
The paper explains why hyperparameters become easier to tune as models scale: as scale increases, the hyperparameter loss surface becomes lower dimensional.
The effective number of hyperparameters (γ), which measures the intrinsic dimension of the loss surface at the optimum, drops to 1 as models scale.
This reduction is driven primarily by increasing parameters rather than data: Parameters, more than data, reduce the effective number of hyperparameters.
Despite the existence of scaling laws at small scales, the authors caution that extrapolation hits statistical limitations.
They show that while scaling laws extrapolate well near the data, extrapolation magnifies small differences due to sampling error,
particularly when estimating the irreducible error term. Therefore, a holistic approach is required
rather than pure quantitative extrapolation.
The paper synthesizes its insights with recent literature to propose a new methodology for model-centric research. This methodology relies on three facts: (1) with the pretraining data held fixed, capabilities depend on pretraining loss alone
(perplexity–capability correspondence); (2) with rigorous hyperparameter tuning, scaling laws for pretraining extend to tiny scales
; and (3) as models scale, they grow far less sensitive to their hyperparameters.
The approach involves tune and measure at the small scale, understand how pretraining loss changes, then carry the winner up.
The authors demonstrate this methodology on a classic question: where to place normalization layers in the transformer architecture.
They compare pre-norm and post-norm transformers and recover the large-scale result that pre-normalization works better as models grow in size.
Through a sequence of diagnostics—checking the noisy quadratic limit, examining hyperparameter sensitivity across scales, verifying perplexity–capability correspondence, and validating scaling laws—they conclude that pre-norm scales better than post-norm near the data
and that post-norm was harder to tune at every turn.
The paper concludes that small-scale experiments can deliver on scaling laws' long-awaited promise
when conducted with the right tools and understanding. The authors emphasize that small-scale experiments save compute, but more than that they make our conclusions more robust.
However, they acknowledge limitations, particularly for data-centric research: Changing the data breaks the perplexity–capability correspondence which pretraining loss requires to proxy for downstream tasks.
Improvements for AI systems
Improvements to AI Systems:
-
Hyperparameter-Aware Scaling Law Predictors: Build scaling law estimators that explicitly model hyperparameter sensitivity as a function of model size. Instead of assuming a single scaling curve, the system would predict a family of curves indexed by hyperparameter configurations, with uncertainty bounds that shrink as model size grows. This allows practitioners to extrapolate performance for a given tuning budget, not just for optimal tuning.
-
Small-Scale Model Selection with Tuned-Frontier Transfer: Develop an automated pipeline that, given a fixed compute budget, performs exhaustive hyperparameter search on tiny models (e.g., 4M–50M parameters) to identify the tuned frontier (Pareto-optimal loss vs. size). The system then uses the learned relationship between hyperparameter sensitivity and scale to transfer the best configuration to larger models, reducing the need for large-scale tuning runs.
-
Effective Hyperparameter Dimensionality (γ) Estimator: Create a diagnostic tool that, during training, estimates the intrinsic dimensionality γ of the loss surface at the optimum. This tool would alert researchers when γ is high (small models, high sensitivity) and recommend more aggressive tuning, or when γ≈1 (large models, low sensitivity) and allow relaxed search. This enables adaptive compute allocation for tuning across scales.
-
Scale-Aware Uncertainty Quantification for Extrapolation: Enhance scaling law extrapolation with statistical confidence intervals that explicitly account for sampling error in the irreducible loss term (γ0). The system would flag when extrapolation beyond the observed range is unreliable (e.g., when the confidence interval explodes) and suggest a
holistic
approach—combining quantitative extrapolation with qualitative diagnostics (e.g., loss–capability correspondence checks) rather than blind extrapolation. -
Perplexity–Capability Correspondence Validator: Build a module that, given a pretraining dataset and a set of downstream tasks, automatically verifies whether perplexity (pretraining loss) is a monotonic proxy for task performance. If the correspondence holds, the system can safely use small-scale perplexity measurements to rank model architectures; if not, it warns that small-scale results may not transfer to capabilities.
-
Architecture Search via Small-Scale Diagnostics: Implement an automated architecture search that, for each candidate (e.g., pre-norm vs. post-norm), runs a battery of small-scale diagnostics: (a) fit a noisy quadratic loss surface to measure curvature and γ, (b) test hyperparameter sensitivity across 2–3 scales, (c) verify perplexity–capability correspondence on a few probe tasks, and (d) fit a scaling law. The system then ranks candidates by scaling slope rather than absolute small-scale loss, enabling correct selection of architectures that improve with scale (e.g., pre-norm).
-
Tuning-Budget-Aware Experiment Planner: A system that, given a compute budget and a target model size, automatically decides how to split resources between (i) exhaustive tuning at small scale, (ii) moderate tuning at medium scale, and (iii) minimal tuning at large scale. It uses the empirical finding that hyperparameter sensitivity fades with scale to minimize wasted compute, while ensuring the small-scale tuned frontier is accurately estimated.
-
Robustness-First Small-Scale Evaluation Framework: A methodology that, for any model-centric research question, mandates: (1) report results on the fully tuned frontier (not just default hyperparameters), (2) report γ values to quantify tuning difficulty, (3) report extrapolation confidence intervals, and (4) validate with at least one diagnostic from the paper's checklist. This framework would be embedded into experiment tracking tools to prevent misleading small-scale conclusions.
What the Improved AI System Can Do:
-
Predict large-model performance from tiny-model runs with quantified uncertainty, reducing compute by 10–100× for architecture and hyperparameter decisions.
-
Automatically identify when a small-scale result is trustworthy (low γ, validated perplexity–capability correspondence, tight extrapolation bounds) versus when it is an artifact of poor tuning.
-
Select architectures (e.g., normalization placement, activation functions) that scale optimally, even when their small-scale performance is inferior, by focusing on scaling slopes and tuned frontiers.
-
Provide researchers with a
tuning difficulty score
per model size, enabling smarter allocation of search budgets. -
Flag data-centric experiments where small-scale transfer is invalid (broken perplexity–capability correspondence), preventing wasted compute.
Abstract
Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver. Instead, researchers have found them unreliable at small scales (starting at 4M parameters) and concluded that sizable models cannot be avoided. We show this is not the case: the confounding factor is hyperparameters. Small models are highly sensitive, but hyperparameter sensitivity fades with scale. This small-scale sensitivity makes scaling laws easy to miss because they only emerge on the fully tuned frontier, and reaching that frontier requires an extensive search far beyond what most ever run. By ablating the basic scaling law recipe, we show well-tuned hyperparameters matter more than any other ingredient. Further, we reveal why those hyperparameters become easier to find: as scale increases, the hyperparameter loss surface becomes lower dimensional. Nevertheless while scaling laws exist in small models, extrapolation hits statistical limitations. A holistic approach is required. Synthesizing our insights with the recent literature, we develop a new methodology for model-centric research and demonstrate it on a question that once took the field years to settle: where to place normalization layers in the transformer architecture. From small-scale experiments, we recover the large scale result: pre-normalization works better as models grow in size. With the right tools and a better understanding, small-scale experiments can deliver on scaling laws' long-awaited promise.
Sources
- PaLM 2 Technical Report
- Adaptive Input Representations for Neural Language Modeling
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- The rising costs of training frontier AI models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- The Llama 3 Herd of Models
- Scaling Laws for Autoregressive Generative Modeling
- Scaling laws for single-agent reinforcement learning
- Scaling Scaling Laws with Board Games
- Scaling Laws for Neural Language Models
- Predictable Scale: Part I, Step Law -- Optimal Hyperparameter Scaling Law in Large Language Model Pretraining
- In-context Learning and Induction Heads
- Scaling Language Models: Methods, Analysis & Insights from Training Gopher
- GLU Variants Improve Transformer
- Configuration-to-Performance Scaling Law with Neural Ansatz
- Random Scaling of Emergent Capabilities
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks