StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models
cs.LG, cs.AI, math.OC
Submitted: 2026-04-16
Updated: 2026-08-25
Code: https://github.com/OptimalScale/LMFlow
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sign-based optimization algorithms, such as SignSGD, have garnered attention for their performance in distributed learning and training large foundation models.
Terminology
Abstract
Sign-based optimization algorithms, such as SignSGD, have garnered attention for their performance in distributed learning and training large foundation models. Despite their empirical superiority, SignSGD is known to diverge on non-smooth objectives, which are ubiquitous due to ReLUs, max-pools, and mixture-of-experts. To overcome this limitation, we propose StoSignSGD, an algorithm that injects structural stochasticity into the sign operator while maintaining an unbiased update step. In the regime of (online) convex optimization, StoSignSGD rigorously resolves the non-convergence issues of SignSGD, achieving a sharp convergence rate matching the lower bound. For the more challenging non-convex non-smooth optimization, we introduce generalized stationary measures that encompass prior definitions, proving that StoSignSGD improves upon the best-known complexity bounds by dimensional factors. Empirically, StoSignSGD is stable and efficient across diverse large language model (LLM) training regimes. In aggressive low-precision pretraining, which spans both FP8 and the far more demanding FP4 regime where AdamW fails catastrophically, StoSignSGD stays stable and consistently performs the best. It attains a 1.44x to 2.14x speedup over established baselines under FP8. Under 4-bit precision, it improves downstream accuracy on the largest OLMo2-370M model by 1.13 points over the strongest stable baseline, and this advantage grows as both the model size and the data scale up. When fine-tuning 7B LLMs on mathematical reasoning tasks, StoSignSGD also delivers clear gains over both AdamW and SignSGD. Finally, to explain why it works, we develop a sign conversion framework that turns any general optimizer into its unbiased, sign-based counterpart. Using this framework, we decompose the core components of StoSignSGD and run a comprehensive ablation study to validate our design choices.
Sources
- Pretraining Large Language Models with NVFP4
- GPT-4 Technical Report
- On the Convergence of SGD with Biased Gradients
- Group Distributionally Robust Optimization with Flexible Sample Queries
- Scalify: scale propagation for efficient low-precision LLM training
- SignSVRG: fixing SignSGD via variance reduction
- Training Verifiers to Solve Math Word Problems
- Convergence Rate Analysis of LION
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Low-Precision Training of Large Language Models: Methods, Challenges, and Opportunities
- Introduction to Online Convex Optimization
- Lecture Notes: Optimization for Machine Learning
- Gaussian Error Linear Units (GELUs)
- Mistral 7B
- Convergence Analysis of the Lion Optimizer in Centralized and Distributed Settings
- Improved Analysis for Sign-based Methods with Momentum Updates
- Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees
- Gemma 3 Technical Report
- To FP8 and Back Again: Quantifying Reduced Precision Effects on LLM Training Stability
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks