Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes

summary

Video file (mp4)

The gist

The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism” [1].

In short

Deep neural networks show periodic loss spikes during long training called Slingshot Mechanism. This paper proves these spikes are caused by floating-point arithmetic limits, specifically when logit differences exceed precision thresholds. This leads to Softmax Collapse and a feedback loop called Numerical Feature Inflation, explaining abnormal parameter growth in late-stage training.

Key concepts

Slingshot Mechanism
Periodic loss spikes during unregularized long training. The paper shows this is not just optimization dynamics but a result of finite-precision arithmetic errors when the difference between correct and incorrect class logits becomes too large for the system to handle.
Softmax Collapse (SC)
A phenomenon during backpropagation where the gradient for the correct class logit is rounded exactly to zero, while gradients for incorrect classes remain non-zero. This breaks standard gradient constraints and initiates a systematic drift in parameter updates.
Numerical Feature Inflation (NFI)
A deterministic feedback loop between global classifier mean and feature mean driven by Softmax Collapse. This loop causes the means of these vectors to grow exponentially after long training, which is what triggers the Slingshot loss spikes.

Terminology used across episodes

This episode discusses

The paper

Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes · Read on arXiv

Tsinghua University · The University of Tokyo

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes".

Jane: The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism”

1: .

Tom: First, who's behind it and why it matters.

Title and authors: Tom: To summarize, this paper argues that those loss spikes we see during long training aren't just random optimization hiccups; they are actually caused by how computers handle numbers when precision gets low during backpropagation.

Jane: It’s about a mechanism called the Slingshot Mechanism where the difference between your correct answer logit and all the other logits gets huge, triggering a rounding error that kills the gradient for your right class but leaves others untouched #pg2.

Lu: Because of that error, it creates this feedback loop called Numerical Feature Inflation where both your main classifier mean and the feature mean start growing exponentially during training #pg3.

Meng: Exponential growth in parameters sounds dangerous for deployment, so what’s the actual fix they are proposing? Is it just forcing everything to use high precision?

Tom: The authors state that casting only the logits and loss computation to float64 is enough to completely remove those spikes because that stops the absorption errors causing the gradient rounding #pg2.

Jane: They also suggest structural fixes, like removing the last layer mean or using smaller learning rates with Adam, which can slow down that Numerical Feature Inflation instability #pg3.

Lu: The paper frames this whole thing as reinterpreting what we call grokking—that sudden jump in accuracy after long training—as a numerical dynamic of finite-precision math #pg2.

Meng: So it shifts the focus away from just adding more regularization to being super careful about the underlying arithmetic when you're running models for a really long time.

Tom: Exactly, and they point out that this mechanism isn't limited to toy models; it shows up in practical large-model training too, which is important for real-world stability #pg1.

Jane: They also mention that while Numerical Feature Inflation explains the rapid growth before a spike, in some real tasks, you might not even see the spike itself, but that drift can still happen and drive parameter norms up quickly #pg3.

Lu: That means we need to be watching those low-precision regimes closely because they seem to become more relevant as we move toward even lower precision formats like BF16 or FP8 #pg2.

Meng: It sounds like a practical warning for anyone deploying large models that long training sessions, you can’t just trust the standard mixed precision settings without checking these dynamics #pg2.

Tom: Absolutely, this work gives us a testable explanation for why late-stage training can suddenly become unstable, which is a big deal for building more robust AI systems #pg1.

Jane: So the main idea here is that we need to treat finite-precision loss computation as a first-order factor in understanding long-term training stability #pg2.

Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking #pg2.

The paper's summary: Tom: So, we’re moving on to how they suggest actually fixing these spikes instead of just identifying the problem—the improvements section of this paper is all about practical ways to tame that numerical inflation we talked about.

Jane: It’s not just about using better hardware anymore; they’re proposing specific changes you can make to your training setup to keep things stable during those long runs #pg3.

Lu: The most direct fix they suggest is using float64 precision for the critical parts of the loss computation, which we already touched on, but they emphasize that this is the key step for stabilization #pg2.

Meng: From an engineer’s view, that makes sense because it directly targets the source—the absorption errors—so if you can eliminate those errors through precision choice, you remove one major source of drift #pg2.

Tom: Exactly, but they also talk about other ways to manage the dynamics, like restoring that zero-sum constraint or even removing parts of the final classifier mean during training to suppress that Numerical Feature Inflation feedback loop #pg3.

Jane: They’re suggesting a couple of structural adjustments, like centering the global feature mean right before the final layer to stop that principal drift component from growing exponentially #pg3.

Lu: That leads into something interesting with the learning rate; they mention tuning the Adam epsilon parameter, and they show that increasing it can reduce how often those spikes happen by lowering that effective learning rate bound #pg2.

Meng: So if I’m running a large model, I can start experimenting with those hyperparameter tweaks to see if adjusting things like epsilon or batch normalization right before the classifier helps keep the parameters from spiraling out of control #pg2.

Tom: That’s the practical application, but they also give us some architectural context; for models like ResNet18, they suggest that because of how their dynamics work, you might need to make sure your epsilon change isn't too big compared to the weight and feature mean growth itself #pg2.

Jane: They also acknowledge a limitation here: this whole mechanism is tied specifically to finite precision arithmetic errors in the mantissa, so if you move into even lower precision formats like BF16 or FP8, those spikes could become a critical instability source in aggressive quantization regimes #pg3.

Lu: That’s a good catch; it means that as we push for more efficient hardware, we need to be extra cautious because the margin required to trigger those absorption errors shrinks dramatically #pg3.

Meng: So the paper is giving us a roadmap: use float64 for critical steps, try structural means adjustments like Batch Normalization, and be aware that lower precision might introduce new kinds of numerical instability #pg2.

Tom: It sounds like a really clear set of action items for anyone training deep neural networks in the long term #pg1.

Jane: It really highlights that when training gets long, we have to think about the math behind the gradients as much as the model structure itself to ensure stability #pg2.

The paper's improvements: Tom: So we’ve covered how finite precision math can trigger a whole feedback loop called Numerical Feature Inflation in deep learning models, and how that leads to those sudden loss spikes we call Slingshot Mechanism.

Jane: Basically, the core idea is that these spikes aren't just random optimization noise; they are a direct consequence of the way computers handle gradients when training gets really deep and confident #pg1.

Lu: The implication for me is huge because it suggests that what we see as grokking under no explicit regularization might actually be a numerical artifact of the system hitting its floating-point limits #pg2.

Meng: And practically, it changes how we think about stability in large models; we can’t just rely on standard mixed precision settings if those low-precision dynamics are causing parameter growth #pg2.

Lalam: If AI can start exhibiting these kinds of numerical instabilities during long training, it means our systems need to be far more robust and predictable for real-world applications where continuous learning is essential #pg1.

Tom: The paper gives us solid ways to mitigate this, like forcing float64 on the loss computation or using Batch Normalization before the final layer to kill that drift component #pg2.

Jane: It really emphasizes that we have tools to counteract this instability, not just observe it happening and wonder why #pg2.

Lu: It opens up a whole new area of research looking at how numerical dynamics interact with the geometry of the loss landscape in ways we haven't explored before #pg2.

Meng: I’m going to keep watching those practical intervention suggestions, like removing the last-layer mean, to see if those changes actually stabilize things in a real deployment setting #pg3.

Lalam: For culture and development, understanding these deep numerical behaviors means we can build AI systems that are fundamentally more reliable and less prone to sudden, unpredictable failures during operation #pg1.

Tom: So it’s a big deal because it gives us a testable explanation for abnormal parameter growth in late-stage training without needing complex new theoretical frameworks immediately #pg1.

Jane: It shifts our view from just looking at the model's structure to paying close attention to the underlying arithmetic governing its training behavior #pg2.

Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking #pg2.

Conclusion: Tom: So we've been diving into "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes." The main takeaway is that those weird loss spikes during long training aren't just random optimization hiccups; they are caused by how computers handle numbers when precision gets low during backpropagation.

Jane: That’s right, it’s a mechanism called the Slingshot Mechanism where if the gap between your correct answer logit and all the others gets too big, you get a rounding error that kills the gradient for your right class but leaves others untouched.

Lu: And because of that rounding error, it creates this feedback loop called Numerical Feature Inflation where both your main classifier mean and the feature mean start growing exponentially during training.

Meng: Exponential growth in parameters sounds scary for deployment, so what’s the actual fix they are proposing? Is it just forcing everything to use high precision?

Tom: The authors say casting only the logits and loss computation to float64 is enough to completely remove those spikes because that stops the absorption errors that cause that gradient rounding.

Jane: They also suggest structural fixes, like removing the last layer mean or using smaller learning rates with Adam, which can slow down that NFI instability and stop things from growing too fast.

Lu: They also touched on Adam optimizers, saying increasing the epsilon parameter can help reduce how often those spikes happen because it lowers the upper bound of the effective learning rate.

Meng: That makes sense for implementation. So, if we're using mixed precision training, this suggests that maybe even fine-tuning those settings could be a lever to control these kinds of instabilities before they get out of hand.

Tom: Exactly. And they point out that some architectures, like ResNet18, might behave differently because the mechanism seems to require a change in epsilon that is smaller than the overall dynamics of the weights and feature means.

Jane: They also mentioned a limitation for this analysis: while NFI explains the rapid growth before a spike, they noted that in more practical tasks, you might not even see the spike itself, but that drift can still happen and drive parameter norms up quickly.

Lu: That means we need to be watching those low-precision regimes closely because they seem to become more relevant as we move toward even lower precision formats like BF16 or FP8.

Meng: I think we need to keep an eye on these low-precision regimes because they might become more prevalent as we move to even lower precision formats like BF16 or FP8.

Tom: And it’s clear that the paper "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes" is important because it provides a testable explanation for why late-stage training can suddenly become unstable.

Jane: It gives us a concrete reason to investigate these loss spikes, moving them out of the realm of just being intrinsic optimization dynamics.

Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking.

Meng: I'll be checking out those practical intervention suggestions, like removing the last-layer mean, to see if it actually stabilizes things in a real deployment setting.

Lalam: For culture and development, understanding these deep numerical behaviors means we can build AI systems that are fundamentally more reliable and less prone to sudden, unpredictable failures during operation.

Tom: So it’s a big deal because it gives us a testable explanation for abnormal parameter growth in late-stage training without needing complex new theoretical frameworks immediately.

Jane: It shifts our view from just looking at the model's structure to paying close attention to the underlying arithmetic governing its training behavior.

Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking.

Meng: I'm ready for whatever comes next on the show.

More episodes

← Home