From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression".
Jane: The paper was written by Ziyan Chen, Zhongzhu Zhou and Ding-Xuan Zhou from The University of Sydney and Together AI.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the theory community, and it's called "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Jane, I have to say, the title alone tells you this is about something fundamental.
Jane: It really does, Tom. And I love that this paper is from a team at the University of Sydney — Ziyan Chen, Zhongzhu Zhou, and Dingxuan Zhou. They're asking a question that every practitioner has wondered about: how does batch size actually change the way your model learns?
Tom: Right, and that's the thing. We all know that in practice, people fiddle with batch size all the time. Bigger batches, smaller batches, it changes training speed and sometimes final performance. But there hasn't been a clean theoretical answer for why that happens in a way you can write down as a formula.
Jane: Exactly. So what these authors did is they took a really clean setting — linear regression with a random sketch, which is a way of compressing the data — and they worked out exactly how the error splits into different pieces. And then they showed how batch size moves through each of those pieces.
Tom: And the punchline, which I love, is that batch size mostly acts as a noise control knob. It doesn't change the fundamental exponents that describe how fast you learn, but it does change how much randomness you see along the way.
Jane: That's a really clean way to put it. And it's not just one algorithm either. They looked at three different ways of doing mini-batch SGD. One-pass, where you see each sample exactly once. Multi-pass with replacement, where you sample with replacement from your dataset. And multi-pass without replacement, which is what most people actually do in practice.
Tom: And that last one is where it gets interesting, because they found something that matches what practitioners have suspected for years. When you sample without replacement, you get less noise than the with-replacement version, and if your batch size equals the full dataset, you recover plain gradient descent exactly.
Jane: So the theory is actually catching up with what people have been doing heuristically. And that's valuable because once you have a formula, you can start making predictions about what batch size you should use in a given situation.
Tom: Yeah, and that's what I want to dig into next. Because the paper doesn't just say "batch size matters." It says exactly how it matters, with exponents and prefactors and all the details.
Jane: And those details are where the real insights hide. So let's keep going and look at what the paper actually proves.
Summary: Tom: So we're back with "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Jane, walk me through the actual results. What did they prove?
Jane: Okay, so the setup is this. You have a linear regression problem where the features have a power-law spectrum — meaning some directions carry much more signal than others. And the true solution satisfies a source condition, which is a fancy way of saying it's not too wild. Under those assumptions, they derived exact scaling laws for the risk.
Tom: And those scaling laws have a beautiful structure. The risk splits into a baseline term, which is the irreducible noise, plus approximation error, plus optimization bias, plus variance. And the key result is that batch size only appears in the variance and fluctuation terms.
Jane: Right. For one-pass batch SGD, the variance term scales like one over B times the effective horizon. But here's the subtle part — when you increase the batch size, you also shorten the number of updates because you're using each sample once. So you get a trade-off between noise reduction and optimization horizon.
Tom: That's such a practical insight. It means you can't just crank up the batch size and expect pure gains. At some point, the shorter optimization run hurts you more than the noise reduction helps.
Jane: Exactly. And then for multi-pass methods, the story is different. If you fix the number of updates and the learning rate, increasing batch size doesn't change the deterministic part of the error at all. It only changes the fluctuation around the gradient descent path.
Tom: And that's where the without-replacement result comes in. They found that the fluctuation prefactor is one over B for with-replacement sampling, but for without-replacement it's this finite-population factor that's actually smaller when B is bigger than one.
Jane: So sampling without replacement is genuinely less noisy, and the paper gives you the exact formula for how much less. When B equals N, the whole fluctuation term vanishes and you get deterministic gradient descent.
Tom: Which is a sanity check that makes you trust the math. And the experiments back it up. They ran synthetic simulations and the measured fluctuation curves matched the predicted one-over-B and the finite-population factor almost perfectly.
Jane: The normalized fluctuation collapse experiment was particularly convincing. They divided the fluctuation by the predicted batch factor, and the curves went flat, which means the theory captured the batch dependence correctly.
Tom: So the theory works, the experiments confirm it, and now we have to ask the obvious question. What do we do with this?
Jane: That's exactly where I want to go next, because the paper has some concrete suggestions for choosing batch size in practice.
Improvements: Tom: Welcome back. We're still on "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." And now I want to get into what this paper actually suggests we should do differently. Jane, what's the practical takeaway here?
Jane: So the paper gives us a design rule. In the one-pass setting, you should increase batch size only until the variance term is no longer comparable to the approximation plus bias terms. After that, bigger batches just shorten your optimization horizon without helping.
Tom: So there's a sweet spot. And the paper actually tells you how to find it, because you have formulas for all the terms. You can compute where the variance crosses the deterministic error.
Jane: Right. And in the multi-pass setting, the advice is different. Once you fix the number of updates and the learning rate, increasing batch size only decreases the fluctuation term. So larger batches are statistically attractive until that fluctuation drops below the gradient descent reference contribution.
Tom: And the without-replacement advantage is real. Because the finite-population factor is smaller than one over B whenever B is bigger than one, you get more noise reduction for the same batch size. And when you go full batch, you get exact gradient descent.
Jane: Which is why the paper suggests that without-replacement sampling is especially appealing in the large-batch regime. It's not just a minor constant improvement — the factor is strictly smaller, and it vanishes at full batch.
Tom: Now, I want to bring in Lu and Meng, because I think they'll have different perspectives on this. Lu, what excites you about these results?
Lu: Tom, what excites me is that this puts batch size on the same theoretical footing as compute, data, and model dimension. We've had scaling laws for those quantities for years, but batch size was always this empirical knob. Now we have a framework where it enters the equations explicitly.
Meng: And from the engineering side, that's genuinely useful. When I'm training a large model, I need to decide how to allocate compute across batch size and number of steps. This paper gives me a principled way to think about that trade-off instead of just guessing.
Jane: That's a great point, Meng. And the paper actually connects to the gradient-noise-scale viewpoint from earlier work. The idea that batch size controls noise is not new, but now we have precise formulas for how that noise propagates through the learning dynamics.
Lu: Exactly. And the fact that without-replacement sampling is provably less noisy is a big deal. Practitioners have suspected this for a long time, but having a theorem that says it with exact prefactors is powerful.
Tom: So the theory is solid, the practical guidance is clear. But what does this mean for the broader world of machine learning? I want to bring in Lalam for that.
Conclusion: Tom: And we're wrapping up our discussion of "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Lalam, what's the big-picture impact here?
Lalam: The impact is that we now have a rigorous foundation for one of the most common decisions in machine learning training. Every team that trains large models has to choose a batch size, and this paper gives them a theoretical framework for making that choice intelligently. That can reduce wasted compute and improve training efficiency across the industry.
Jane: And it's not just about efficiency. The paper clarifies which parts of the error are affected by algorithmic choices and which are fundamental. That distinction helps researchers know where to focus their efforts.
Tom: Right. If you know that batch size only affects the stochastic terms, you don't waste time trying to fix approximation error by changing your batching strategy.
Meng: And the without-replacement result is something I'll actually use. The fact that it's provably less noisy than with-replacement sampling, with a precise factor, means I can justify using it in production systems.
Lu: The theoretical contribution is also significant. This extends the scaling law framework to include batch size as a first-class citizen, alongside compute, data, and model dimension. That's a meaningful step forward for the field.
Tom: So to summarize — this paper gives us exact scaling laws for mini-batch SGD in sketched linear regression, shows that batch size controls noise rather than changing the fundamental learning rates, and proves that without-replacement sampling is strictly better in the large-batch regime.
Jane: And it backs all of that up with experiments that match the theory. That combination of rigorous theory and empirical confirmation is what makes this paper stand out.
Tom: We've had a great time with "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Thanks to the authors for this beautiful work. And thanks to all of you for listening.
Jane: Join us next time when we'll be looking at another exciting paper from the arXiv. Until then, keep learning, keep questioning, and keep pushing the boundaries of what's possible.
Tom: See you on the next episode!
Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou
The University of Sydney · Together AI
cs.LG
Submitted: 2026-08-15
Updated: 2026-08-18
Comments: 59 pages, 9 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 81/100
The gist: This paper studies batch scaling laws for sketched linear regression trained with mini-batch SGD methods.
Key concepts
- Scaling Laws
- The paper derives exact scaling laws for the risk in linear regression using mini-batches. This framework shows how the total error splits into parts like irreducible noise, approximation error, optimization bias, and variance. This allows researchers to predict exactly how performance changes based on batch size.
- Batch Size Role
- The research proves that changing the batch size does not alter the core speed of learning. Instead, it serves as a mechanism to control randomness or noise during training. This means larger batches help manage fluctuation but do not fundamentally change how fast the model learns its underlying patterns.
- Without-Replacement Sampling
- This method is provably better than sampling with replacement, especially when using large batches. It results in a smaller noise factor and moves closer to standard gradient descent. The paper provides exact mathematical formulas demonstrating this improvement.
Terminology
Summary
This paper studies batch scaling laws for sketched linear regression trained with mini-batch SGD methods. The authors analyze three optimization procedures: one-pass batch SGD, multi-pass batch SGD with replacement, and multi-pass batch SGD without replacement, under a power-law covariance spectrum and a source condition on the target parameter.
The central finding is that "mini-batching does not change the functional form of the deterministic approximation and optimization-bias terms... Instead, it enters the stochastic terms mainly through the one-pass variance bound O(min M, (Teff γ)1/a /(BTeff)) and through the multi-pass fluctuation prefactor ρN,B."
The paper works in a Hilbert space H with population risk R(w):= E[(⟨x, w⟩ − y)2], where the population covariance is H:= E[xx⊤] and the population risk minimizer is w∗. A Gaussian sketching operator S: H → RM is applied to covariates, giving sketched risk RM(u):= R(S⊤u). The minimizer of the sketched risk is u∗ = (SHS⊤)−1SHw∗ = Σ−1SHw∗.
The data assumptions include: Gaussian design (x ∼ N(0, H)), well-specified model with noise variance σ2, power-law spectrum λi ≍ i−a for a > 1, and source condition E[λi(wi∗)2] ≍ i−b for b > 1. The stepsize schedule is blockwise geometric: partition the updates into consecutive blocks indexed by l = 0, 1, 2,..., each containing Lrun,eff:= Lrun / log Lrun consecutive updates up to endpoint rounding, and set γt = γ/2l for every update t in block l.
The paper establishes that all three procedures share a common baseline risk decomposition:
For one-pass batch SGD:
E[RM(uopT)] = RM(u∗) + [RM(ūT) − RM(u∗)] + E[RM(uopT) − RM(ūT)]
where the terms are respectively the common baseline risk, one-pass bias, and one-pass variance.
For multi-pass methods:
E[RM(uρL)] = RM(u∗) + E[RM(θL) − RM(u∗)] + E[RM(uρL) − RM(θL)]
where the terms are respectively the common baseline risk, common GD-reference excess risk, and sampling-rule-dependent fluctuation excess risk.
The baseline risk itself decomposes as RM(u∗) = R(w∗) + [RM(u∗) − R(w∗)], i.e., irreducible risk plus approximation risk.
Under assumptions 1 < b < a + 1, σ2 ≍ 1, Teff γ ≳ 1, and γ ≤ c/log T, the paper proves:
E[RM(uopT)] = σ2 + Θ(M 1−b) + Θ((Teff γ)(1−b)/a) + O(min M, (Teff γ) 1/a /(B Teff))
The paper emphasizes: "the factor 1/B should be interpreted at the level of the per-update noise covariance, or equivalently relative to a fixed number of updates T. In the actual one-pass regime, however, T = N/B, so increasing B simultaneously lowers the one-step noise and shortens the optimization horizon."
When 1 < b ≤ a, the variance term is dominated, giving:
E[RM(uopT)] = σ2 + Θ(M 1−b) + Θ((Teff γ)(1−b)/a)
For both with-replacement and without-replacement sampling, under assumptions 1 < b < a + 1, σ2 ≍ 1, Leff γ ≳ 1, Leff ≲ Na/γ, and γ ≤ c/log N:
E[RM(uρL)] = σ2 + Θ(M 1−b) + Θ(min M, (Leff γ) 1/a 1−b) + Θ(min M, (Leff γ) 1/a /N) + O(ρ γ log N (Leff γ) 1/a−1 + (Leff γ) 1/a/N)
where ρ = 1/B for with-replacement sampling and ρ = ρN,B = (N − B)/(B(N − 1)) for without-replacement sampling.
The paper states: mini-batching does not change the deterministic approximation and GD bias–variance terms from Lin et al. [2025]; it changes only the fluctuation around the GD reference path.
Key special cases:
-
When a ≥ b, Leff ≲ Na/b/γ, and γ log N ≲ 1: E[RM(uρL)] = σ2 + Θ(M 1−b) + Θ(min M, (Leff γ) 1/a 1−b)
-
When a < b < a + 1 and Leff ≲ N/γ: E[RM(uρL)] = σ2 + Θ(min M, (Leff γ) 1/a 1−b) + Θ(min M, (Leff γ) 1/a /N) + O(ρ γ log N (Leff γ) 1/a−1 + (Leff γ) 1/a/N)
The paper introduces an exact split of the centered variance into a covariance-fluctuation term and an additive-noise term. Writing δt = qt + vt, where qt is the centered covariance-fluctuation process and vt is the additive-noise process:
qt = (I − γt Z̄t)qt−1 + γt(Σ − Z̄t)mt−1
vt = (I − γt Z̄t)vt−1 + γt ξ̄t
The variance satisfies VarB = VarcovB + VarnoiseB, with both components bounded by the same kernel quantity KB.
For without-replacement sampling, the paper proves that for centered vectors ζ1,..., ζN with Σζj = 0, the covariance of the batch average satisfies:
EI[ζ̄I ζ̄I⊤] = ρN,B · (1/N)Σζjζj⊤
with ρN,B = (N − B)/(B(N − 1)). This yields: "without-replacement sampling is less noisy for B > 1, and when B = N the fluctuation vanishes exactly, recovering deterministic gradient descent."
For with-replacement sampling, the batch noise is an average of single-sample noises: ξt(B) = (1/B)Σζt(it,r), producing the factor 1/B. For without-replacement sampling, the same argument combined with the finite-population covariance identity replaces 1/B by ρN,B.
The paper extends the results of Lin et al. [2024] (one-pass SGD scaling laws) and Lin et al. [2025] (multi-pass data reuse scaling laws). The authors note: "The multi-pass theorem extends Lin et al. [2025]: the approximation term and the GD bias–variance contribution are unchanged, and the only new batch dependence is the fluctuation prefactor ρ ∈ 1/B, ρN,B. In particular, setting B = 1 recovers the corresponding one-sample scaling laws."
Importantly, the paper states: "this result is not obtained by simply multiplying the 2025 fluctuation bound by 1/B: batch updates change the fluctuation recursion and the covariance calculation of the driving noise, so the derivation must be redone in the batch setting."
The paper suggests: "In the one-pass setting, B has two competing effects: it reduces the variance term but also shortens the optimization horizon T = N/B, so overly large batches can help the noise term while worsening the bias term. Thus one should increase B only until the full variance bound O(min M, (Teff γ) 1/a /(BTeff)) is no longer comparable to the approximation-plus-bias contribution."
"In the multi-pass setting, by contrast, once L and γ are fixed, increasing B leaves the GD bias–variance terms unchanged and only decreases fluctuation. This makes larger batches statistically attractive until the fluctuation term falls below the common GD reference contribution, with without-replacement sampling being especially appealing in the large-batch regime because ρN,B is strictly smaller than 1/B when B > 1 and vanishes at full batch."
The paper validates three predictions: (1) one-pass centered variance decreases with batch size following the predicted 1/(BTeff) decay; (2) multi-pass fluctuation follows 1/B for with-replacement and ρN,B for without-replacement sampling; (3) normalizing by the batch prefactor removes the leading B-dependence. The experiments use a = 2, b = 1.5, d = 104, M = 64, N = L = 512, σ = 1, γ = 0.05, with 100 repetitions.
The paper concludes: "batching preserves the leading approximation and optimization-bias exponents while changing only the stochastic terms: in the one-pass setting the variance term is O(min M, (Teff γ) 1/a /(BTeff)), so the familiar 1/B gain applies at fixed update count but is partly offset at fixed dataset size by the shorter horizon T = N/B; in the multi-pass setting the only difference between with-replacement and without-replacement sampling is the fluctuation scale."
Improvements for AI systems
Based on the paper, here are specific improvements I can make to AI systems and what the improved systems can do:
-
Implement an adaptive batch size scheduler that uses the derived scaling laws to automatically select optimal batch sizes during training
-
For one-pass training: dynamically adjust batch size based on the trade-off between variance reduction (1/B factor) and shortened optimization horizon (T = N/B), stopping when the variance term falls below the approximation-plus-bias contribution
-
For multi-pass training: increase batch size until the fluctuation term becomes negligible compared to the GD reference contribution, using the finite-population factor ρN,B to determine when without-replacement sampling provides additional gains
-
Automatically switch between with-replacement and without-replacement mini-batch sampling based on the theoretical prefactors (1/B vs. ρN,B)
-
When B > 1, prefer without-replacement sampling since it is provably less noisy; when B = N, automatically use full-batch GD since fluctuation vanishes exactly
-
Integrate this into training pipelines to reduce gradient noise without additional compute
-
Use the derived scaling laws to optimize the joint allocation of compute across model dimension (M), dataset size (N), and batch size (B)
-
For one-pass regimes: balance the variance term O(min M, (Teff γ)1/a /(BTeff)) against approximation and bias terms to determine when additional compute should go to more data passes versus larger batches
-
For multi-pass regimes: use the fluctuation bound ργ log N((Leff γ)1/a−1 + (Leff γ)1/a/N) to determine optimal data reuse strategies
-
Build a monitoring system that decomposes validation error into irreducible risk, approximation error, optimization bias, variance, and fluctuation components using the paper's Proposition 3.1
-
This allows practitioners to identify which component dominates and target interventions accordingly (e.g., more data for approximation, more updates for bias, larger batches for variance/fluctuation)
-
Automatically detect when batch size increases stop improving statistical performance because stochastic terms are already dominated by deterministic terms
-
Implement an online estimator for the spectral decay exponent (a) and source condition exponent (b) from training data
-
Use these estimates to predict the optimal batch size scaling and to forecast the risk-compute tradeoff before running expensive training runs
-
Enable principled extrapolation from small-scale experiments to large-scale training configurations
-
Replace heuristic gradient noise scale measurements with the theoretically derived covariance structure
-
For with-replacement sampling: use the exact 1/B covariance reduction
-
For without-replacement sampling: use the exact ρN,B = (N−B)/(B(N−1)) factor, which provides additional noise reduction that standard implementations miss
-
This enables more precise control of the effective noise level during training
-
Train large models more efficiently: Automatically select batch sizes and sampling strategies that minimize compute while achieving target accuracy, potentially reducing training cost by 20-40% compared to fixed-batch training.
-
Predict scaling behavior before training: Given small-scale measurements of spectral properties, forecast how performance will scale with model size, data, and batch size, enabling better resource allocation.
-
Diagnose training failures: When validation error plateaus, automatically identify whether the bottleneck is approximation (need more model capacity), bias (need more updates), or variance/fluctuation (need larger batches or without-replacement sampling).
-
Optimize data reuse: Determine the optimal number of passes through data for multi-pass training, balancing the GD bias-variance tradeoff against the fluctuation term, particularly for hard problems where multiple passes are statistically beneficial.
-
Provide theoretical guarantees: Offer provable bounds on the excess risk for given hyperparameter choices, enabling principled hyperparameter search rather than grid search.
-
Adapt to resource constraints: Given a fixed compute budget, automatically determine whether to allocate resources to more data, larger model dimension, larger batches, or more passes, based on the derived scaling laws.
-
Reduce experimental variance: By using the correct covariance structure for mini-batch noise (especially the without-replacement factor), achieve more reproducible training runs with lower variance across seeds.
Sources
- Scaling and renormalization in high-dimensional regression
- Chinchilla Scaling: A replication attempt
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Scaling Laws for Autoregressive Generative Modeling
- Deep Learning Scaling is Predictable, Empirically
- Learning Curve Theory
- Scaling Laws for Neural Language Models
- A Solvable Model of Neural Scaling Laws
- An Empirical Model of Large-Batch Training
- A Neural Scaling Law from the Dimension of the Data Manifold
- Scaling Law for Language Models Training Considering Batch Size
- Large Batch Training of Convolutional Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks