Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes".
Jane: The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism”
1: .
Tom: First, who's behind it and why it matters.
Title and authors: Tom: To summarize, this paper argues that those loss spikes we see during long training aren't just random optimization hiccups; they are actually caused by how computers handle numbers when precision gets low during backpropagation.
Jane: It’s about a mechanism called the Slingshot Mechanism where the difference between your correct answer logit and all the other logits gets huge, triggering a rounding error that kills the gradient for your right class but leaves others untouched #pg2.
Lu: Because of that error, it creates this feedback loop called Numerical Feature Inflation where both your main classifier mean and the feature mean start growing exponentially during training #pg3.
Meng: Exponential growth in parameters sounds dangerous for deployment, so what’s the actual fix they are proposing? Is it just forcing everything to use high precision?
Tom: The authors state that casting only the logits and loss computation to float64 is enough to completely remove those spikes because that stops the absorption errors causing the gradient rounding #pg2.
Jane: They also suggest structural fixes, like removing the last layer mean or using smaller learning rates with Adam, which can slow down that Numerical Feature Inflation instability #pg3.
Lu: The paper frames this whole thing as reinterpreting what we call grokking—that sudden jump in accuracy after long training—as a numerical dynamic of finite-precision math #pg2.
Meng: So it shifts the focus away from just adding more regularization to being super careful about the underlying arithmetic when you're running models for a really long time.
Tom: Exactly, and they point out that this mechanism isn't limited to toy models; it shows up in practical large-model training too, which is important for real-world stability #pg1.
Jane: They also mention that while Numerical Feature Inflation explains the rapid growth before a spike, in some real tasks, you might not even see the spike itself, but that drift can still happen and drive parameter norms up quickly #pg3.
Lu: That means we need to be watching those low-precision regimes closely because they seem to become more relevant as we move toward even lower precision formats like BF16 or FP8 #pg2.
Meng: It sounds like a practical warning for anyone deploying large models that long training sessions, you can’t just trust the standard mixed precision settings without checking these dynamics #pg2.
Tom: Absolutely, this work gives us a testable explanation for why late-stage training can suddenly become unstable, which is a big deal for building more robust AI systems #pg1.
Jane: So the main idea here is that we need to treat finite-precision loss computation as a first-order factor in understanding long-term training stability #pg2.
Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking #pg2.
The paper's summary: Tom: So, we’re moving on to how they suggest actually fixing these spikes instead of just identifying the problem—the improvements section of this paper is all about practical ways to tame that numerical inflation we talked about.
Jane: It’s not just about using better hardware anymore; they’re proposing specific changes you can make to your training setup to keep things stable during those long runs #pg3.
Lu: The most direct fix they suggest is using float64 precision for the critical parts of the loss computation, which we already touched on, but they emphasize that this is the key step for stabilization #pg2.
Meng: From an engineer’s view, that makes sense because it directly targets the source—the absorption errors—so if you can eliminate those errors through precision choice, you remove one major source of drift #pg2.
Tom: Exactly, but they also talk about other ways to manage the dynamics, like restoring that zero-sum constraint or even removing parts of the final classifier mean during training to suppress that Numerical Feature Inflation feedback loop #pg3.
Jane: They’re suggesting a couple of structural adjustments, like centering the global feature mean right before the final layer to stop that principal drift component from growing exponentially #pg3.
Lu: That leads into something interesting with the learning rate; they mention tuning the Adam epsilon parameter, and they show that increasing it can reduce how often those spikes happen by lowering that effective learning rate bound #pg2.
Meng: So if I’m running a large model, I can start experimenting with those hyperparameter tweaks to see if adjusting things like epsilon or batch normalization right before the classifier helps keep the parameters from spiraling out of control #pg2.
Tom: That’s the practical application, but they also give us some architectural context; for models like ResNet18, they suggest that because of how their dynamics work, you might need to make sure your epsilon change isn't too big compared to the weight and feature mean growth itself #pg2.
Jane: They also acknowledge a limitation here: this whole mechanism is tied specifically to finite precision arithmetic errors in the mantissa, so if you move into even lower precision formats like BF16 or FP8, those spikes could become a critical instability source in aggressive quantization regimes #pg3.
Lu: That’s a good catch; it means that as we push for more efficient hardware, we need to be extra cautious because the margin required to trigger those absorption errors shrinks dramatically #pg3.
Meng: So the paper is giving us a roadmap: use float64 for critical steps, try structural means adjustments like Batch Normalization, and be aware that lower precision might introduce new kinds of numerical instability #pg2.
Tom: It sounds like a really clear set of action items for anyone training deep neural networks in the long term #pg1.
Jane: It really highlights that when training gets long, we have to think about the math behind the gradients as much as the model structure itself to ensure stability #pg2.
The paper's improvements: Tom: So we’ve covered how finite precision math can trigger a whole feedback loop called Numerical Feature Inflation in deep learning models, and how that leads to those sudden loss spikes we call Slingshot Mechanism.
Jane: Basically, the core idea is that these spikes aren't just random optimization noise; they are a direct consequence of the way computers handle gradients when training gets really deep and confident #pg1.
Lu: The implication for me is huge because it suggests that what we see as grokking under no explicit regularization might actually be a numerical artifact of the system hitting its floating-point limits #pg2.
Meng: And practically, it changes how we think about stability in large models; we can’t just rely on standard mixed precision settings if those low-precision dynamics are causing parameter growth #pg2.
Lalam: If AI can start exhibiting these kinds of numerical instabilities during long training, it means our systems need to be far more robust and predictable for real-world applications where continuous learning is essential #pg1.
Tom: The paper gives us solid ways to mitigate this, like forcing float64 on the loss computation or using Batch Normalization before the final layer to kill that drift component #pg2.
Jane: It really emphasizes that we have tools to counteract this instability, not just observe it happening and wonder why #pg2.
Lu: It opens up a whole new area of research looking at how numerical dynamics interact with the geometry of the loss landscape in ways we haven't explored before #pg2.
Meng: I’m going to keep watching those practical intervention suggestions, like removing the last-layer mean, to see if those changes actually stabilize things in a real deployment setting #pg3.
Lalam: For culture and development, understanding these deep numerical behaviors means we can build AI systems that are fundamentally more reliable and less prone to sudden, unpredictable failures during operation #pg1.
Tom: So it’s a big deal because it gives us a testable explanation for abnormal parameter growth in late-stage training without needing complex new theoretical frameworks immediately #pg1.
Jane: It shifts our view from just looking at the model's structure to paying close attention to the underlying arithmetic governing its training behavior #pg2.
Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking #pg2.
Conclusion: Tom: So we've been diving into "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes." The main takeaway is that those weird loss spikes during long training aren't just random optimization hiccups; they are caused by how computers handle numbers when precision gets low during backpropagation.
Jane: That’s right, it’s a mechanism called the Slingshot Mechanism where if the gap between your correct answer logit and all the others gets too big, you get a rounding error that kills the gradient for your right class but leaves others untouched.
Lu: And because of that rounding error, it creates this feedback loop called Numerical Feature Inflation where both your main classifier mean and the feature mean start growing exponentially during training.
Meng: Exponential growth in parameters sounds scary for deployment, so what’s the actual fix they are proposing? Is it just forcing everything to use high precision?
Tom: The authors say casting only the logits and loss computation to float64 is enough to completely remove those spikes because that stops the absorption errors that cause that gradient rounding.
Jane: They also suggest structural fixes, like removing the last layer mean or using smaller learning rates with Adam, which can slow down that NFI instability and stop things from growing too fast.
Lu: They also touched on Adam optimizers, saying increasing the epsilon parameter can help reduce how often those spikes happen because it lowers the upper bound of the effective learning rate.
Meng: That makes sense for implementation. So, if we're using mixed precision training, this suggests that maybe even fine-tuning those settings could be a lever to control these kinds of instabilities before they get out of hand.
Tom: Exactly. And they point out that some architectures, like ResNet18, might behave differently because the mechanism seems to require a change in epsilon that is smaller than the overall dynamics of the weights and feature means.
Jane: They also mentioned a limitation for this analysis: while NFI explains the rapid growth before a spike, they noted that in more practical tasks, you might not even see the spike itself, but that drift can still happen and drive parameter norms up quickly.
Lu: That means we need to be watching those low-precision regimes closely because they seem to become more relevant as we move toward even lower precision formats like BF16 or FP8.
Meng: I think we need to keep an eye on these low-precision regimes because they might become more prevalent as we move to even lower precision formats like BF16 or FP8.
Tom: And it’s clear that the paper "Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes" is important because it provides a testable explanation for why late-stage training can suddenly become unstable.
Jane: It gives us a concrete reason to investigate these loss spikes, moving them out of the realm of just being intrinsic optimization dynamics.
Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking.
Meng: I'll be checking out those practical intervention suggestions, like removing the last-layer mean, to see if it actually stabilizes things in a real deployment setting.
Lalam: For culture and development, understanding these deep numerical behaviors means we can build AI systems that are fundamentally more reliable and less prone to sudden, unpredictable failures during operation.
Tom: So it’s a big deal because it gives us a testable explanation for abnormal parameter growth in late-stage training without needing complex new theoretical frameworks immediately.
Jane: It shifts our view from just looking at the model's structure to paying close attention to the underlying arithmetic governing its training behavior.
Lu: We should keep looking at how these numerical features inflate and what those anti-parallel alignments between the weights and feature means are doing there because it’s a deep dive into the math behind grokking.
Meng: I'm ready for whatever comes next on the show.
Tsinghua University · The University of Tokyo
cs.LG, cs.CL, math.OC, stat.ML
Submitted: 2026-05-07
Updated: 2026-10-07
Comments: 29 pages, 13 figures; accepted to NeurIPS 2026
Code: https://github.com/google-deepmind/nanodo
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 82/100
The gist: The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism” [1].
Key concepts
- Slingshot Mechanism
- Periodic loss spikes during unregularized long training. The paper shows this is not just optimization dynamics but a result of finite-precision arithmetic errors when the difference between correct and incorrect class logits becomes too large for the system to handle.
- Softmax Collapse (SC)
- A phenomenon during backpropagation where the gradient for the correct class logit is rounded exactly to zero, while gradients for incorrect classes remain non-zero. This breaks standard gradient constraints and initiates a systematic drift in parameter updates.
- Numerical Feature Inflation (NFI)
- A deterministic feedback loop between global classifier mean and feature mean driven by Softmax Collapse. This loop causes the means of these vectors to grow exponentially after long training, which is what triggers the Slingshot loss spikes.
Terminology
Summary
The gist Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism” [1].
How it works
-
The Slingshot Mechanism is proven to be a result of floating-point arithmetic precision limits, specifically triggered when the difference between the correct-class logit and the other logits may exceed the absorption-error threshold [3].
-
This leads to Softmax Collapse (SC), where during backpropagation,
the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero
[3]. -
SC breaks the zero-sum constraint of gradients across classes, which introduces a systematic drift in the parameter update of the classifier layer [3].
-
This drift forms a positive feedback loop with the feature, causing
the global classifier mean and the global feature mean to grow exponentially
in a process called Numerical Feature Inflation (N FI) [3]. -
N FI explains how Slingshot loss spikes are triggered by driving the system into a regime where
the sample-level absorption condition becomes fragile
[3]. -
The exponential growth is formalized by Theorem 3.7, showing that the norms of the global weight mean WG and feature mean µG exhibit exponential growth after long-term training [3].
Numerical Feature Inflation (N FI)
The mechanism of Numerical Feature Inflation is a deterministic feedback loop between the global classifier mean and the global feature mean that drives parameter and logit growth after long-term unregularized training [3].
-
The weight mean update is given by E[∆WG] = −ηϵ/(KµG) [3].
-
This drift necessitates a redefinition of geometric alignment, where the weight drift WG is in the direction of −µG [3].
-
The feature vector receives a non-zero gradient component parallel to WG, ProjWG(∇hL) = ϵWG [3].
-
The coupled dynamics lead to an asymptotic relationship where
cos(W(t)G, µ(t)G) → −1
and the norms grow exponentially [3].
Mitigation and Context
-
Casting only the logits/loss computation to float64 is sufficient to remove the spikes, showing that
the instability originates from the loss computation
[3]. -
Practical interventions include restoring the zero-sum constraint or removing the last-layer mean, which suppresses N FI-induced instability and parameter growth [3].
-
Increasing the Adam optimizer’s ε parameter mitigates instability:
while ε = 10−6 reduces the frequency of spikes, setting ε = 10−5 completely eliminates them
[8]. -
Applying Batch Normalization (BN) directly before the final classifier successfully eliminates slingshots by removing
the principal drift component
[8]. -
Label Smoothing (LS) sets the target probability to y < 1, which prevents the infinite growth of correct-class logits but introduces a new class of instabilities independent of numerical precision [8].
Empirical Findings
-
Slingshot spikes are observed across various architectures and datasets, with most models exhibiting them except ResNet18 [3].
-
The occurrence of Slingshots is independent of dataset size, confirming that
as long as the model possesses sufficient capacity to drive the residual probability mass ϵ below the floating-point absorption threshold, the N FI feedback loop will be triggered
[14]. -
Smaller learning rates are significantly more prone to triggering Slingshot instabilities, implying that
a larger learning rate provides a crucial form of implicit regularization
[15]. -
The mechanism is driven purely by absorption errors in the mantissa, independent of exponent-based underflow, as demonstrated by experiments where logits were clamped [14].
-
The model exhibits a connection to grokking, as each loss spike is accompanied by a stepwise increase in test accuracy, which may help the model reach
flatter and more generalizable solutions
[3].
Conclusion
Finite-precision loss computation should be treated as a first-order factor in the analysis of long-term training stability because SC and N FI are not limited to toy models but also appear in practical large-model training [10]. The work provides a testable explanation for abnormal parameter growth and logit divergence in late-stage training by reinterpreting Slingshot as a numerical dynamic of finite-precision training [3].
--- Page 1 ---
Grokking or Glitching? How Low-Precision Drives Slingshot Loss Spikes Liu Hanqing Jianjun Cao Yuanze Li Zijian Zhou 1Tsinghua University 2The University of Tokyo Abstract Deep neural networks exhibit periodic loss spikes during unregularized longterm training, a phenomenon known as the “Slingshot Mechanism” [1]. Existing work usually attributes this to intrinsic optimization dynamics, but its triggering mechanism remains unclear. This paper proves that this phenomenon is a result of floating-point arithmetic precision limits. As training enters a high-confidence stage, the difference between the correct-class logit and the other logits may exceed the absorption-error threshold. Then during backpropagation, the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero. This breaks the zero-sum constraint of gradients across classes and introduces a systematic drift in the parameter update of the classifier layer. We prove that this drift forms a positive feedback loop with the feature, causing the global classifier mean and the global feature mean to grow exponentially
[3]. We call this mechanism Numerical Feature Inflation (N FI). This mechanism explains the rapid norm growth before a Slingshot spike, the subsequent reappearance of gradients, and the resulting loss spike
[3]. We further show that N FI is not equivalent to an observed loss spike: in more practical tasks, partial absorption may not produce visible spikes, but it can still break the zero-sum constraint and drive rapid growth of parameter norms [3]. Our results reinterpret Slingshot as a numerical dynamic of finite-precision training, and provide a testable explanation for abnormal parameter growth and logit divergence in late-stage training [3]. 1 Introduction Loss spikes are a persistent puzzle in neural network training, creating difficulties for both theoretical understanding and practical stability. One representative example is the Slingshot Mechanism, first observed in the study of grokking under no explicit regularization. Grokking refers to the phenomenon where neural networks achieve sudden generalization long after reaching perfect training accuracy [2]. In such settings, training is often accompanied by periodic instabilities: the norm of the last-layer parameters grows rapidly, often close to exponentially, and is followed by an abrupt training loss spike
[2]. Existing work has mainly interpreted Slingshot as an intrinsic optimization phenomenon. For example, Thilak et al. [1] related it to the Edge of Stability (EOS) [3], where the optimizer periodically crosses stability boundaries. Nanda et al. [4] suggested that the effect may arise from the interaction between gradients of different scales and adaptive optimizer dynamics. In contrast, we show that Slingshot is not primarily caused by the intrinsic optimization dynamics of the model. Instead, it is triggered by finite-precision arithmetic in the computation of the cross-entropy (CE) loss [3]. After long training, the model enters a high-confidence regime: it reaches perfect training accuracy, and for each training sample, the correct-class logit becomes much larger than the other logits
[3]. Preprints arXiv:2605.06152v3 [cs.LG] 26 May 2026
--- Page 2 ---
Epoch Train Loss (Log) Float32 Float64 Mixed Precision (a) 0 50 100 150 200 Epoch N o r m o f G G o.2 G o.4 G o.6 G o.8 1.0 None SC sample Count Norm of G (b) 25 Epoch Train Loss (Log) Float32 Float64 Mixed Precision (a) 0 50 100 150 200 Epoch N o r m o f G o.2 G o.4 G o.6 G o.8 N o r m of G (b) C sin e Similarity Between WG and WG Cosine Similarity Target Logit Average Target Logit Figure 1: Precision-induced N FI dynamics (a) Slingshot loss spikes disappear when training is performed in float64. Casting only the logits/loss computation to float64 is also sufficient to remove the spikes, showing that the instability originates from the loss computation
[3]. (b) Before most samples enter Softmax Collapse, the global feature mean grows slowly
[3]. Once most samples collapse, ∥µG∥ enters a rapid-growth phase
[3]. (c) The cosine similarity between WG and µG approaches −1. Under this anti-parallel alignment, even the correct-class logits become deeply negative [3]. Prieto et al. [5] showed that when the gap between the largest logit zm and the other logits exceeds a threshold determined by floating-point precision, absorption error causes the computed result to differ from the exact real-number value
[3]. As a result, during backpropagation, the gradient of the loss with respect to the correct-class logit, ∂L/∂zm, is rounded exactly to zero
[3]. They call this phenomenon Softmax Collapse (SC) [3]. We show that SC has a second effect that directly drives Slingshot.
Improvements for AI systems
- Bold header: Numerical Precision Mitigation in Loss Computation
Implement loss computation using float64 precision for critical stages of training to eliminate Slingshot spikes, as Casting only the logits/loss computation to float64 is sufficient to remove the spikes.
This stabilizes training by preventing absorption errors that cause the gradient of the correct class is rounded exactly to zero,
which drives instability.
- Bold header: Dynamic Feature Mean Stabilization
Employ a mechanism equivalent to removing or centering the global feature mean, such as applying Batch Normalization immediately before the final classifier, to suppress Numerical Feature Inflation (N FI). This addresses the principal drift component
by removing µG during training,
which substantially slows down the late-stage parameter growth.
- Bold header: Adaptive Learning Rate Control
Utilize a smaller Adam optimizer learning rate or explicitly increase the Adam epsilon parameter to mitigate instability, as increasing it causes the system to lower the upper bound of the effective learning rate (η/εAdam),
thereby preventing the amplification of the re-emerging gradient signals.
- Bold header: Architectural-Specific Stability Checks
For models like ResNet18, monitor training dynamics closely; if instability is avoided, it suggests that N FI mechanism requires the change of ϵ to be negligible compared with the dynamics of W and µ,
allowing for more aggressive low-precision training in those architectures.
- Bold header: Precision-Aware Regularization
In low-precision paradigms (BF16, FP8), account for the shrinking margin required to trigger absorption errors, as the margin required to trigger absorption errors shrinks dramatically,
suggesting that spikes induced by N FI could be a critical source of instability in these aggressive quantization regimes.
Abstract
Deep neural networks exhibit periodic loss spikes during unregularized long-term training, a phenomenon known as the "Slingshot Mechanism." Existing work usually attributes this to intrinsic optimization dynamics, but its triggering mechanism remains unclear. This paper proves that this phenomenon is a result of floating-point arithmetic precision limits. As training enters a high-confidence stage, the difference between the correct-class logit and the other logits may exceed the absorption-error threshold. Then during backpropagation, the gradient of the correct class is rounded exactly to zero, while the gradients of the incorrect classes remain nonzero. This breaks the zero-sum constraint of gradients across classes and introduces a systematic drift in the parameter update of the classifier layer. We prove that this drift forms a positive feedback loop with the feature, causing the global classifier mean and the global feature mean to grow exponentially. We call this mechanism Numerical Feature Inflation (NFI). This mechanism explains the rapid norm growth before a Slingshot spike, the subsequent reappearance of gradients, and the resulting loss spike. We further show that NFI is not equivalent to an observed loss spike: in more practical tasks, partial absorption may not produce visible spikes, but it can still break the zero-sum constraint and drive rapid growth of parameter norms. Our results reinterpret Slingshot as a numerical dynamic of finite-precision training, and provide a testable explanation for abnormal parameter growth and logit divergence in late-stage training.
Sources
- The Slingshot Mechanism: An Empirical Study of Adaptive Optimizers and the Grokking Phenomenon
- Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
- Grokking modular arithmetic
- Progress Measures for Grokking on Real-world Tasks
- Grokking and Generalization Collapse: Insights from \texttt{HTSR} theory
- Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- A Theory on Adam Instability in Large-Scale Machine Learning
- Adaptive Preconditioners Trigger Loss Spikes in Adam
- Output Embedding Centering for Stable LLM Pretraining
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks