From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression

summary

Video file (mp4)

The gist

This paper studies batch scaling laws for sketched linear regression trained with mini-batch SGD methods.

In short

This episode discusses a paper detailing scaling laws for mini-batch Stochastic Gradient Descent (SGD) in linear regression. The authors developed exact formulas showing how batch size influences error components. Key findings include specific design rules for choosing batch size and proving that sampling without replacement is more efficient than with replacement.

Key concepts

Scaling Laws
The paper derives exact scaling laws for the risk in linear regression using mini-batches. This framework shows how the total error splits into parts like irreducible noise, approximation error, optimization bias, and variance. This allows researchers to predict exactly how performance changes based on batch size.
Batch Size Role
The research proves that changing the batch size does not alter the core speed of learning. Instead, it serves as a mechanism to control randomness or noise during training. This means larger batches help manage fluctuation but do not fundamentally change how fast the model learns its underlying patterns.
Without-Replacement Sampling
This method is provably better than sampling with replacement, especially when using large batches. It results in a smaller noise factor and moves closer to standard gradient descent. The paper provides exact mathematical formulas demonstrating this improvement.

Terminology used across episodes

This episode discusses

The paper

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression · Read on arXiv

Ziyan Chen, Zhongzhu Zhou, Ding-Xuan Zhou

The University of Sydney · Together AI

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression".

Jane: The paper was written by Ziyan Chen, Zhongzhu Zhou and Ding-Xuan Zhou from The University of Sydney and Together AI.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back to the show, everyone. Today we're looking at a paper that's been making waves in the theory community, and it's called "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Jane, I have to say, the title alone tells you this is about something fundamental.

Jane: It really does, Tom. And I love that this paper is from a team at the University of Sydney — Ziyan Chen, Zhongzhu Zhou, and Dingxuan Zhou. They're asking a question that every practitioner has wondered about: how does batch size actually change the way your model learns?

Tom: Right, and that's the thing. We all know that in practice, people fiddle with batch size all the time. Bigger batches, smaller batches, it changes training speed and sometimes final performance. But there hasn't been a clean theoretical answer for why that happens in a way you can write down as a formula.

Jane: Exactly. So what these authors did is they took a really clean setting — linear regression with a random sketch, which is a way of compressing the data — and they worked out exactly how the error splits into different pieces. And then they showed how batch size moves through each of those pieces.

Tom: And the punchline, which I love, is that batch size mostly acts as a noise control knob. It doesn't change the fundamental exponents that describe how fast you learn, but it does change how much randomness you see along the way.

Jane: That's a really clean way to put it. And it's not just one algorithm either. They looked at three different ways of doing mini-batch SGD. One-pass, where you see each sample exactly once. Multi-pass with replacement, where you sample with replacement from your dataset. And multi-pass without replacement, which is what most people actually do in practice.

Tom: And that last one is where it gets interesting, because they found something that matches what practitioners have suspected for years. When you sample without replacement, you get less noise than the with-replacement version, and if your batch size equals the full dataset, you recover plain gradient descent exactly.

Jane: So the theory is actually catching up with what people have been doing heuristically. And that's valuable because once you have a formula, you can start making predictions about what batch size you should use in a given situation.

Tom: Yeah, and that's what I want to dig into next. Because the paper doesn't just say "batch size matters." It says exactly how it matters, with exponents and prefactors and all the details.

Jane: And those details are where the real insights hide. So let's keep going and look at what the paper actually proves.

Summary: Tom: So we're back with "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Jane, walk me through the actual results. What did they prove?

Jane: Okay, so the setup is this. You have a linear regression problem where the features have a power-law spectrum — meaning some directions carry much more signal than others. And the true solution satisfies a source condition, which is a fancy way of saying it's not too wild. Under those assumptions, they derived exact scaling laws for the risk.

Tom: And those scaling laws have a beautiful structure. The risk splits into a baseline term, which is the irreducible noise, plus approximation error, plus optimization bias, plus variance. And the key result is that batch size only appears in the variance and fluctuation terms.

Jane: Right. For one-pass batch SGD, the variance term scales like one over B times the effective horizon. But here's the subtle part — when you increase the batch size, you also shorten the number of updates because you're using each sample once. So you get a trade-off between noise reduction and optimization horizon.

Tom: That's such a practical insight. It means you can't just crank up the batch size and expect pure gains. At some point, the shorter optimization run hurts you more than the noise reduction helps.

Jane: Exactly. And then for multi-pass methods, the story is different. If you fix the number of updates and the learning rate, increasing batch size doesn't change the deterministic part of the error at all. It only changes the fluctuation around the gradient descent path.

Tom: And that's where the without-replacement result comes in. They found that the fluctuation prefactor is one over B for with-replacement sampling, but for without-replacement it's this finite-population factor that's actually smaller when B is bigger than one.

Jane: So sampling without replacement is genuinely less noisy, and the paper gives you the exact formula for how much less. When B equals N, the whole fluctuation term vanishes and you get deterministic gradient descent.

Tom: Which is a sanity check that makes you trust the math. And the experiments back it up. They ran synthetic simulations and the measured fluctuation curves matched the predicted one-over-B and the finite-population factor almost perfectly.

Jane: The normalized fluctuation collapse experiment was particularly convincing. They divided the fluctuation by the predicted batch factor, and the curves went flat, which means the theory captured the batch dependence correctly.

Tom: So the theory works, the experiments confirm it, and now we have to ask the obvious question. What do we do with this?

Jane: That's exactly where I want to go next, because the paper has some concrete suggestions for choosing batch size in practice.

Improvements: Tom: Welcome back. We're still on "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." And now I want to get into what this paper actually suggests we should do differently. Jane, what's the practical takeaway here?

Jane: So the paper gives us a design rule. In the one-pass setting, you should increase batch size only until the variance term is no longer comparable to the approximation plus bias terms. After that, bigger batches just shorten your optimization horizon without helping.

Tom: So there's a sweet spot. And the paper actually tells you how to find it, because you have formulas for all the terms. You can compute where the variance crosses the deterministic error.

Jane: Right. And in the multi-pass setting, the advice is different. Once you fix the number of updates and the learning rate, increasing batch size only decreases the fluctuation term. So larger batches are statistically attractive until that fluctuation drops below the gradient descent reference contribution.

Tom: And the without-replacement advantage is real. Because the finite-population factor is smaller than one over B whenever B is bigger than one, you get more noise reduction for the same batch size. And when you go full batch, you get exact gradient descent.

Jane: Which is why the paper suggests that without-replacement sampling is especially appealing in the large-batch regime. It's not just a minor constant improvement — the factor is strictly smaller, and it vanishes at full batch.

Tom: Now, I want to bring in Lu and Meng, because I think they'll have different perspectives on this. Lu, what excites you about these results?

Lu: Tom, what excites me is that this puts batch size on the same theoretical footing as compute, data, and model dimension. We've had scaling laws for those quantities for years, but batch size was always this empirical knob. Now we have a framework where it enters the equations explicitly.

Meng: And from the engineering side, that's genuinely useful. When I'm training a large model, I need to decide how to allocate compute across batch size and number of steps. This paper gives me a principled way to think about that trade-off instead of just guessing.

Jane: That's a great point, Meng. And the paper actually connects to the gradient-noise-scale viewpoint from earlier work. The idea that batch size controls noise is not new, but now we have precise formulas for how that noise propagates through the learning dynamics.

Lu: Exactly. And the fact that without-replacement sampling is provably less noisy is a big deal. Practitioners have suspected this for a long time, but having a theorem that says it with exact prefactors is powerful.

Tom: So the theory is solid, the practical guidance is clear. But what does this mean for the broader world of machine learning? I want to bring in Lalam for that.

Conclusion: Tom: And we're wrapping up our discussion of "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Lalam, what's the big-picture impact here?

Lalam: The impact is that we now have a rigorous foundation for one of the most common decisions in machine learning training. Every team that trains large models has to choose a batch size, and this paper gives them a theoretical framework for making that choice intelligently. That can reduce wasted compute and improve training efficiency across the industry.

Jane: And it's not just about efficiency. The paper clarifies which parts of the error are affected by algorithmic choices and which are fundamental. That distinction helps researchers know where to focus their efforts.

Tom: Right. If you know that batch size only affects the stochastic terms, you don't waste time trying to fix approximation error by changing your batching strategy.

Meng: And the without-replacement result is something I'll actually use. The fact that it's provably less noisy than with-replacement sampling, with a precise factor, means I can justify using it in production systems.

Lu: The theoretical contribution is also significant. This extends the scaling law framework to include batch size as a first-class citizen, alongside compute, data, and model dimension. That's a meaningful step forward for the field.

Tom: So to summarize — this paper gives us exact scaling laws for mini-batch SGD in sketched linear regression, shows that batch size controls noise rather than changing the fundamental learning rates, and proves that without-replacement sampling is strictly better in the large-batch regime.

Jane: And it backs all of that up with experiments that match the theory. That combination of rigorous theory and empirical confirmation is what makes this paper stand out.

Tom: We've had a great time with "From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression." Thanks to the authors for this beautiful work. And thanks to all of you for listening.

Jane: Join us next time when we'll be looking at another exciting paper from the arXiv. Until then, keep learning, keep questioning, and keep pushing the boundaries of what's possible.

Tom: See you on the next episode!

More episodes

← Home