Adaptive Optimization via Momentum on Variance-Normalized Gradients
summary
The gist
variance-based normalization and applying momentum after normalization.
In short
The episode discusses MVN-Grad, a new optimizer from UT Austin and Google Research, based on 'Adaptive Optimization via Momentum on Variance-Normalized Gradients.' It addresses instability in standard optimizers like Adam by replacing the second moment with variance for normalization and applying momentum after normalization. The authors prove this method is more stable, avoids loss spikes, and matches or beats current state-of-the-art models.
Key concepts
- Optimizer
- In neural network training, an optimizer decides how to adjust the millions of parameters (or 'dials') of a model. It uses the gradients it calculates to determine the direction and magnitude of each step, aiming to improve task performance.
- Adam
- Adam is a widely used optimizer that adapts step size for each parameter individually. It uses a running average of squared gradients (the second moment) to normalize updates, making it efficient but susceptible to issues like loss spikes and sudden jumps.
- Variance Normalization
- This technique estimates the variance—the spread or noise—of the gradients rather than using the raw second moment. By isolating this noise, it helps prevent known failure modes in optimization where gradient information is lost.
- Momentum-After
- This refers to applying momentum after normalizing the current gradient. This ordering prevents old, stale momentum from being amplified by sudden changes in current noise, a problem that occurs when using standard methods like Adam.
Terminology used across episodes
This episode discusses
- Adaptive Optimization via Momentum on Variance-Normalized Gradients · Paper Radio
- Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks
- On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization
- Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights
- Adam: A Method for Stochastic Optimization
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-training
- On the Variance of the Adaptive Learning Rate and Beyond
- Decoupled Weight Decay Regularization
- Adaptive Gradient Methods with Dynamic Bound of Learning Rate
- On the Convergence of Adam and Beyond
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutes
- EAdam Optimizer: How epsilon Impact Adam
- AdaShift: Decorrelation and Convergence of Adaptive Learning Rate Methods
- LaProp: Separating Momentum and Adaptivity in Adam
The paper
Adaptive Optimization via Momentum on Variance-Normalized Gradients · Read on arXiv
Francisco Patitucci, Aryan Mokhtari
University of Texas at Austin · Google Research
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Adaptive Optimization via Momentum on Variance-Normalized Gradients".
Jane: The paper was written by Francisco Patitucci and Aryan Mokhtari from University of Texas at Austin and Google Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a fresh arXiv paper called "Adaptive Optimization via Momentum on Variance-Normalized Gradients." Jane, what's your first read on that title?
Jane: Tom, I love it because it tells you exactly what the fix is. We've got this family of optimizers—Adam being the famous one—that powers basically all of deep learning. And this paper is saying, hey, we can do better by changing two specific things about how Adam works.
Tom: And for our listeners who might not live and breathe optimization, can you break down what an optimizer even does?
Jane: Sure. When you train a neural network, you're trying to adjust millions of little dials to make the network better at its task. The optimizer is the thing that decides how much to turn each dial, and in which direction, based on the gradients it sees. Adam has been the workhorse because it adapts the step size for each dial individually.
Tom: Right, and the paper's title points at two tweaks. One is about using variance instead of the raw second moment, and the other is about when you apply momentum. I've got to say, as someone who's trained models, I've seen Adam do some weird things—loss spikes, sudden jumps.
Jane: Exactly. And the authors—Francisco Patitucci and Aryan Mokhtari from UT Austin and Google Research—they've identified why those weird things happen. The way Adam is structured, there's this coupling between your momentum buffer and your current gradient noise. It's like carrying a heavy backpack while someone randomly changes how heavy the ground is beneath you.
Tom: That's a great way to put it. And their fix, MVN-Grad, applies momentum after normalizing the gradients instead of before. We'll get into the details in a bit, but the headline is that they've got both theory and experiments showing this is more stable.
Jane: And the experiments are on real benchmarks—CIFAR-one hundred image classification and GPT-style language modeling. So this isn't just a toy result. They're matching or beating Adam, AdaBelief, and LaProp across the board.
Tom: So the title is really promising two things: stability and better performance. And from what I've seen in the abstract, they deliver on both. But I'm curious about the variance part—what does that actually change in practice?
Jane: That's the part that gets me excited. Adam normalizes by the average of squared gradients, which mixes together the signal and the noise. MVN-Grad instead estimates the variance, which isolates the noise. In low-noise situations, that means you don't collapse into just taking the sign of the gradient, which is a known failure mode.
Tom: So you preserve more information about how big the gradient actually is. That sounds like it could make training faster in those regimes where the model is already fairly confident.
Jane: Exactly. And they prove that in the paper—there's a theorem showing that second-moment methods can slow down by a factor that depends on the dimension, while their method doesn't. That's a big deal for high-dimensional problems like language models.
Tom: Well, I'm hooked. Let's get into the meat of how they actually changed the update rule.
Summary: Tom: So we're back, still on "Adaptive Optimization via Momentum on Variance-Normalized Gradients." Jane, you gave us the elevator pitch. Now let's talk about what the paper actually does step by step.
Jane: So the algorithm is pretty simple. At each step, you compute your gradient. Then you keep a running average of the gradient itself—that's your estimate of the mean. Then you compute the squared difference between the current gradient and that mean, and you keep a running average of that too. That's your variance estimate.
Tom: And that variance estimate is what you divide by, instead of Adam's second moment.
Jane: Right. And then—this is the key ordering change—you normalize the current gradient by that variance, and only after that do you apply momentum. In Adam, you apply momentum first and then normalize. That ordering difference is what breaks the coupling between stale momentum and fresh noise.
Tom: I remember you used that backpack analogy earlier. Can you take it a step further for our listeners?
Jane: Sure. Imagine you're walking and you've got momentum from your past steps. With Adam, you're dividing that whole momentum—including everything from the past—by a number that's based only on what's happening right now. If the current gradient is unexpectedly small, that denominator shrinks, and suddenly your old momentum gets amplified. That's how you get those spikes.
Tom: And MVN-Grad avoids that because you normalize the current gradient first, and then the momentum is built from already-normalized pieces. So a spike gets clipped at the source.
Jane: Exactly. And they prove this formally. Theorem three point one shows that their method has strictly smaller one-step conditional variance than AdaBelief, which is the variance-based version of Adam. The math is clean—the variance gap is proportional to the square of the momentum times the variance of the inverse normalizer.
Tom: That's the kind of result that makes me trust the empirical results more. It's not just "we tried it and it worked"—there's a structural reason why it should be more stable.
Jane: And they also have a robustness result. If you hit a single giant gradient spike—say a bad batch of data—their method's update stays bounded. Adam can actually amplify that spike later, because the momentum remembers it while the normalizer forgets it. There's a specific time when that mismatch peaks, and they show the update can grow with the spike size.
Tom: So Adam stores the spike and replays it. MVN-Grad just clips it and moves on. That's a really concrete failure mode that I've definitely seen in practice.
Jane: Me too. And then there's the third theoretical piece, which is about the variance normalizer itself. In low-noise regimes, Adam's second-moment normalizer makes the update behave like a sign operation—you lose all information about gradient magnitude. Their method keeps that magnitude information, and they prove it avoids a dimension-dependent slowdown.
Tom: So we've got three theoretical results: lower variance, spike robustness, and no sign collapse. And then they back it up with experiments. What did they find?
Jane: On CIFAR-one hundred with ResNet-eighteen they match AdaBelief at batch size one hundred twenty-eight and beat it at batch size one thousand twenty-four. On language modeling, they get the best perplexity on WikiText-one hundred three and the best validation loss on OpenWebText with a 124M parameter GPT-two model.
Tom: And the hyperparameter robustness plots look good too—narrower spreads across the sweep, which means you don't have to tune as carefully.
Jane: Right. That's actually a huge practical benefit. If an optimizer is less sensitive to the beta choices, that saves real time for practitioners.
Tom: So the summary is: a simple change to the update rule, backed by theory, and it works on real benchmarks. What's not to like? Let's dig into the improvements they're claiming in more detail.
Improvements: Tom: Welcome back. We're still on "Adaptive Optimization via Momentum on Variance-Normalized Gradients." Jane, we've covered the basics and the theory. Now let's talk about what this actually improves over the current state of the art.
Jane: So the paper positions itself in a two-by-two design space. One axis is what you normalize by—second moment versus variance. The other axis is when you apply momentum—before or after normalization. Adam is second moment plus momentum-before. AdaBelief is variance plus momentum-before. LaProp is second moment plus momentum-after. And MVN-Grad is variance plus momentum-after.
Tom: So they're filling in the missing corner of the grid.
Jane: Exactly. And the improvement comes from both axes working together. The variance normalizer fixes the sign-collapse problem, and the momentum-after ordering fixes the temporal coupling problem. Previous methods only fixed one of those at a time.
Tom: That's a really clean way to see it. LaProp already did momentum-after, but it kept the second moment. AdaBelief already did variance, but it kept the old ordering. So this paper is the first to combine both fixes.
Jane: Right. And the theory shows why the combination matters. The conditional variance result in Theorem three point one is specifically about the ordering—it shows that momentum-after is strictly better when you're using a variance normalizer. And the sign-collapse result in Theorem three point three is specifically about the normalizer—it shows that variance is strictly better than second moment in low-noise regimes.
Tom: So each fix addresses a different failure mode, and you need both.
Jane: And there's a nice practical detail in the algorithm. They use the same beta for the mean estimator and the momentum buffer, so you don't need an extra hyperparameter. But they note in a remark that you could decouple them if you wanted to, for non-stationary settings.
Tom: I like that they kept the interface simple. One of the reasons Adam is so popular is that it just works with default hyperparameters. Adding more knobs would hurt adoption.
Jane: And the experiments support that. On OpenWebText with the 124M GPT-two model, they fixed the learning rate at 1e-four for all optimizers and only swept the betas. MVN-Grad still came out on top, and it had the narrowest spread across the sweep.
Tom: So it's more robust to hyperparameter choice. That's a real improvement for people who can't afford massive tuning runs.
Jane: And the training curves are smoother too. They show that in the MNIST toy example—fewer spikes, less variance in the loss. That's the kind of thing that makes training runs more predictable, which matters when you're spending GPU-hours.
Tom: Let me ask you about the practical side. Is there any computational overhead?
Jane: That's the best part. It's essentially the same cost as Adam. You're computing the same moving averages, just in a different order and with a different formula for the normalizer. No extra memory, no extra compute per step.
Tom: So it's a drop-in replacement. You could swap it into an existing training pipeline without changing anything else.
Jane: Exactly. And that's why I think this could actually get adopted. It's not a fundamentally different paradigm—it's a smarter version of what people are already using.
Tom: Before we wrap up, I want to bring in the bigger picture. What does this mean for the field?
Jane: I think it means we're getting closer to optimizers that are both fast and stable. The sign-collapse issue has been known for a while, and the temporal coupling issue too. This paper shows that fixing both at once is not just possible—it's actually better than fixing either one alone.
Conclusion: Tom: And that brings us to the end of our discussion on "Adaptive Optimization via Momentum on Variance-Normalized Gradients." Jane, give us the final takeaway.
Jane: So the paper gives us a new optimizer that changes two things about Adam. It normalizes by variance instead of second moment, and it applies momentum after normalization instead of before. Both changes are backed by theory, and the experiments show it matches or beats Adam, AdaBelief, and LaProp on image classification and language modeling.
Tom: And the best part is that it's free—same computational cost, same memory footprint, no extra hyperparameters to tune.
Jane: Right. And the theory gives us confidence that the improvements are structural, not just lucky hyperparameter choices. The variance gap result, the spike robustness result, and the sign-collapse avoidance result all point in the same direction.
Tom: I think the biggest implication is for large-scale training. If you're training a massive language model, stability and hyperparameter robustness are worth a lot. A smoother training curve means fewer restarts, less wasted compute.
Jane: And for the research community, this fills in a corner of the design space that was previously empty. It shows that the ordering of normalization and momentum matters, and that variance-based normalization has real advantages beyond just AdaBelief's empirical success.
Tom: So what's next? Where do we go from here?
Jane: The paper mentions that decoupling the mean estimator and momentum buffer could help in non-stationary settings. That's an obvious next step. Also, the theory assumes some idealized conditions—like the EMA tracking the mean perfectly—so relaxing those assumptions would be valuable.
Tom: And I'd love to see this tested on even larger models. The 124M parameter GPT-two is a good start, but modern models are ten to a hundred times bigger.
Jane: Absolutely. And also on other domains—vision transformers, diffusion models, reinforcement learning. The failure modes they address are pretty general, so the fix should transfer.
Tom: Well, that's our show for today. We've said goodbye to "Adaptive Optimization via Momentum on Variance-Normalized Gradients," and I'm already looking forward to the next paper on our list.
Jane: Thanks for listening, everyone. If you're training models and you're tired of loss spikes, give MVN-Grad a try. It might just save you a few restarts.
Tom: And remember, the paper is on arXiv, so you can read the full details yourself. Until next time, keep optimizing.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language