Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases".
Jane: The paper was written by Jingwen Liu, Ezra Edelman, Surbhi Goel and Bingbin Liu from Columbia University and University of Pennsylvania and Kempner Institute, Harvard University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: We've just touched on how counterintuitive this idea is, but let's talk about what that title really means. The authors are claiming that "Less Data, Faster Training" isn't some temporary anomaly. They say repeating smaller datasets can speed up learning because of these sampling biases.
Jane: The core concept they are introducing is the small-vs-large gap, which is observing a significant performance improvement when training on fewer samples compared to a larger dataset with the same amount of total computational effort. It's not just about having fewer examples; it’s about how many times those examples are repeated during gradient updates.
Lu: The authors confirm this effect across various tasks and architectures, which is really significant because in AI, we usually see things broken down by specific domains or model types. Showing this trend applies to different settings suggests a universal principle of optimization bias at play here.
Meng: I'm particularly interested in the claim that this gap persists across various algorithmic tasks, even those that are highly structured like sparse parity or single-index models. This is reassuring for me because it means the effect isn's limited to just one type of problem we might solve with AI.
Lalam: It suggests that if we structure our training data repetition correctly, we could be achieving a level of learning efficiency that truly redefines how complex tasks are learned by AI systems. This could be a major shift in how people interact with AI tools.
Summary: Tom: So, the paper summarizes these findings and suggests that this speedup is fundamentally driven by sampling biases. They argue it's not just some random statistical fluctuation but a structural property of the data usage itself.
Jane: The paper shows that training on smaller datasets provides favorable optimization biases, which is more pronounced when the dataset size is smaller. Think of it as the model getting a stronger, more targeted signal from a small set of repeated examples compared to the diluted signal from a massive, diverse dataset.
Lu: This aligns with their mathematical formalization in Section four. It shows that training on smaller datasets can actually reduce the total number of steps required for convergence in certain settings, which is a powerful theoretical result.
Meng: I think the practical implication here is that we might be able to stop over-collect and over-train data just to satisfy the old belief that more is better. If this mechanism works as intended, it could allow us to start designing training pipelines with much tighter constraints on data volume.
Lalam: It opens up a way for AI development where resource constraints are naturally managed through an inherent bias in the training process itself, making AI more scalable and less demanding on a global level.
Improvements: Tom: Now, let's talk about the improvements that the paper suggests, which is really where it becomes proactive rather than reactive. The authors emphasize that this small-vs-large gap isn't just a fallback for data scarcity; it can be leveraged as an inductive bias.
Jane: They show that using more repetitions on a smaller dataset acts like an implicit layerwise preconditioner, steering the model toward a more favorable feature learning regime. It’s not just random repetition; it’s purposeful optimization shaping.
Lu: The theoretical results are strong because they identify regimes where existing theories fail to explain this phenomenon at all, particularly when we look at full-batch gradient updates, which is a massive departure from previous work focusing only on stochastic gradients.
Meng: From my perspective, the practical improvement lies in finding ways to implement these interventions—like adjusting initialization or layer-wise learning rates—to make the model more robust to hyperparameter choices. The data use strategy itself becomes a design choice for robustness.
Lalam: By suggesting that we can proactively use this bias, it implies that AI could be guided toward solutions that are not just faster but also inherently more stable and predictable in terms lead to better long-term cultural impact.
Conclusion: Tom: So, after all these discussions about the mechanics and potential implications, we've seen a lot of compelling evidence across different tasks and architectures. The paper really characterizes this small-vs-large gap thoroughly.
Jane: It seems like the consensus is that "Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases" offers a very promising path forward where efficiency and effectiveness go hand in hand.
Lu: I think we are looking at a fundamental shift in how we approach optimization, moving beyond just data quantity to understanding the subtle interaction between data structure and the model' is capacity.
Meng: It’s definitely something that needs to be incorporated into our engineering practices, not as an optional trick but as a deliberate strategy for efficiency.
Lalam: I feel this research suggests that AI has learned how to find paths of maximum efficiency when we are allowed to understand and control the subtle biases in its own learning process.
Tom: We've heard from Lu, Meng, and Lalam today on the impact of this work. It’s been a fascinating journey through these concepts.
Jane: Thank you all for joining us on this topic of "Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases."
Tom: We've got some excellent insights to share for our listeners and we look forward to the next paper.
Columbia University · University of Pennsylvania · Kempner Institute, Harvard University
cs.LG, cs.AI
Submitted: 2026-05-19
Updated: 2026-09-04
Comments: ICML 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: The paper investigates the "small-vs-large gap," a counterintuitive phenomenon where training on fewer, repeated samples can lead to reduced training compute for a given model compared to using a
Key concepts
- Small-vs-Large Gap
- This concept observes a significant performance improvement when training on fewer samples compared to a larger dataset. The key is not the sheer number of examples, but how many times those smaller examples are repeated during gradient updates.
- Sampling Biases
- The paper argues that the speedup in learning is fundamentally driven by structural properties of data usage. Small datasets provide favorable optimization biases, giving the model a stronger, more targeted signal than massive, diverse datasets.
- Optimization Bias
- This refers to the inherent bias that allows AI systems to learn efficiently when training data is structured or repeated correctly. It suggests that controlling data repetition can guide the model toward a more favorable feature learning regime.
Terminology
Summary
The paper investigates the small-vs-large gap,
a counterintuitive phenomenon where training on fewer, repeated samples can lead to reduced training compute for a given model compared to using a larger dataset. This effect is observed across various algorithmic tasks, architectures, and optimizers—a finding that cannot be explained by existing theory—and suggests that using smaller datasets with more repetitions can be proactively leveraged as a favorable inductive biases for optimization.
The Scope of the Small-vs-Large Gap
The study confirms that the small-vs-large gap exists across a variety of settings, including different tasks (e.g., sparse parity, single-index model), architectures (MLP and Transformers), and optimizers. The gap is evident in both the number of optimization steps and the overall compute complexity, which depends on both the number of steps and the per-step cost proportional to the batch size.
-
The gap persists even under full-batch gradient updates, implying that explanations based solely on stochastic gradient updates are insufficient.
-
The phenomenon is observed in diverse tasks such, as:
-
(20, 6)-sparse parity (Figure 1 and Figure 2).
-
Single-index model (SIM).
-
In-context linear regression.
-
Modular addition.
How the Mechanism Works: Sampling Biases The primary driver of the small-vs-large gap is identified as sampling biases induced by smaller datasets.
This mechanism modulates the relative magnitude of updates across layers, which in turn helps with feature learning.
-
The theory formalizes this intuition, showing that training on smaller datasets reduces the number of steps required for convergence (Theorem 1).
-
The bias enables a
stronger sampling bias
when the dataset is smaller, making the model more robust to learning rate and initialization choice. -
Empirical evidence shows that during initial training phases,
the weight norm ratio a 2 / W F increases more rapidly when the dataset is smaller.
Theoretical Implications The paper establishes that the small-vs-large gap provides a direct saving in the number of steps (T).
-
The analysis shows that Phase 1, using a randomly sampled dataset of size N, allows for faster growth of the outer weight.
-
The total number of steps is shown to be O(N 1/4 d epsilon (1/delta)).
-
The theory suggests that the acceleration comes from the
anti-concentration
of the empirical moment, which is largely independent of the true label signal.
Empirical Interventions and Evidence The study provides extensive evidence through various interventions:
-
The speedup persists even when training on a small set with random labels, confirming that
sampling bias is the main mechanism.
-
Parameter-wise interventions can substantially reduce or eliminate the gap:
-
Adjusting initialization scales (e.scaling) to close the gap.
-
Applying layer-wise learning rates (eta 1 eta 2) to bridge the small-vs-large gap in MLPs and Transformers, making training
more robust
to hyperparameter choices.
- The authors conclude that
training on a smaller dataset with increased repetitions is not just a fallback strategy under data scarcity, but a source of beneficial optimization inductive biases.
Improvements for AI systems
Based on the rigorous analysis of this scientific paper, here are specific, actionable improvements for AI training systems and what the resulting improved systems can achieve.
Improvement: Instead of relying solely on a large dataset (S) for all training steps, implement a two-phase or multi-phase schedule that proactively leverages sampling biases.
-
Phase 1 (Bootstrap/Shaping): Train on a small, randomly sampled subset (N S) for an initial set of steps (T 1). This phase is designed to exploit the strong initial sampling bias, where the inner layer weights accelerate their growth at a rate proportional to O(N 1/2).
-
Phase 2 (Coverage/Generalization): Transition from the small subset training to optimize on the full population (S).
-
Mechanism: This approach directly addresses the
small-vs-large gap
by using the initial high-bias signal from a small set to rapidly establish a favorable layer norm ratio (e.g., a 2/W squared), which then sustains learning for the remaining steps. -
System Capability: The resulting system will achieve significantly faster convergence in total compute complexity, reducing the required number of optimization steps by an order of O((Nd) 1/4), while avoiding the poor generalization associated with simply running a small dataset to completion.
Improvement: For specific algorithmic tasks (like sparse parity or modular addition), implement a dynamic data reuse policy that prioritizes repeating small, biased subsets during the early stages of training, rather than using fresh large batches.
-
Action: During periods where the feature learning signal is still weak (i.e., initial iterations), the system selects a small subset S i and repeats its sampling process multiple times, allowing the internal variables (q(t)) to build up sufficient magnitude (about (1/N)).
-
Mechanism: This leverages the
sampling bias
mechanism identified in Section 4.2, which modulates the relative update speeds across layers. The repetition acts as an implicit layer-wise preconditioner. -
System Capability: The system will exhibit accelerated feature learning and faster alignment between layers, leading to a more robust and efficient path toward convergence for tasks where strong initial correlation is beneficial.
Improvement: Implement adaptive layer-wise adjustments to mitigate the sensitivity of the model to initialization scales or global learning rates, effectively closing the small-vs-large gap.
-
Action:
-
Initialization Scaling (mu P): Employ specific initial scaling factors (e.g, alpha-scaling or mu P) that adjust the standard deviation of layer weights based on their position in a two-layer network. This compensates for the expected initial under-growth of inner weights when using small datasets.
-
Layer-Wise Learning Rates: Apply distinct learning rates (eta 1 eta 2) to the input and subsequent layers, respectively. This mimics the beneficial relative norm growth observed in Phase 1.
-
Mechanism: These interventions stabilize the training dynamics, making the system inherently more robust to hyperparameter tuning and allowing it to capitalize on small-set training biases without catastrophic divergence or premature overfitting.
-
System Capability: The model becomes less sensitive to initial configuration (e.g, GELU vs. Sigmoid initialization) and can reliably achieve performance gains when the data budget is constrained, ensuring that resource limitation does not necessitate a suboptimal training regimen.
Improvement: Develop an architecture-aware decision tree to determine if data repetition is beneficial based on the task structure and model depth/width.
-
Action: Implement checks for certain classes of problems (e.g., structured reasoning tasks like parity or SIM). If the problem exhibits a strong bias that can be amplified by sampling (as defined by O(N 1/2) growth), flag the system to prioritize small-set repetition over full-batch updates.
-
Mechanism: The system identifies that the benefit is not universal. It specifically avoids applying this strategy to non-convex optimization problems like standard linear regression, where it has been shown no such gap exists.
-
System Capability: The AI system can intelligently allocate compute resources, maximizing efficiency by leveraging data repetition only when it provides a genuine inductive bias advantage, thereby avoiding wasted computation on tasks that require standard statistical approaches.
Abstract
This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed across algorithmic tasks, architectures and optimizers and cannot be explained using prior theory. We argue that the speedup comes from appropriate layer-wise growth enabled by sampling biases, which is more pronounced when the dataset size is smaller. We provide both theoretical analysis and empirical evidence from various interventions. Our results suggest that using a smaller dataset with more repetitions is not just a fallback strategy under data scarcity, but can be proactively leveraged as a favorable inductive biases for optimization, particularly in reasoning tasks.
Sources
- Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
- Emergent properties with repeated examples
- Query-Key Normalization for Transformers
- Scaling Laws and Interpretability of Learning from Repeated Data
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
- Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
- Improved Scaling Laws in Linear Regression via Data Reuse
- Small-scale proxies for large-scale Transformer training instabilities
- Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
- Why Does Multi-Epoch Training Help?
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
- Feature Learning in Infinite-Width Neural Networks
- Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
- A Spectral Condition for Feature Learning
- Stabilizing Transformer Training by Preventing Attention Entropy Collapse
- The emergence of sparse attention: impact of data distribution and benefits of repetition
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks