Training nGPT
cs.LG, cs.AI
Submitted: 2026-08-02
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere.
Terminology
Abstract
The normalized Transformer (nGPT) realizes hyperspherical representation learning by constraining model parameter vectors and activation vectors to the unit hypersphere. In this paper, we describe a practical training recipe for nGPT and evaluate it on modern hybrid Mamba-2--Transformer Mixture-of-Experts (MoE) models. The recipe introduces Logit Gradient Preconditioning, Logarithmic Learning Rate Decay, GatedAdamW, angular update control, and optional exploration mechanisms. Compared with an unnormalized model of the same hybrid MoE architecture trained with AdamW, the 30B-total-parameter nGPT model reaches the same validation loss using approximately half as many training tokens. The recipe scales across the models considered, which contain up to 30B total parameters.
Sources
- Nemotron 3 Nano: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Improving Deep Learning Optimization through Constrained Parameter Regularization
- MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
- SPAM: Spike-Aware Adam with Momentum Reset for Stable LLM Training
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
- Weight Norm Control
- SGDR: Stochastic Gradient Descent with Warm Restarts
- An Empirical Study of Mamba-based Language Models
- A Walk with SGD
- Spherical Latent Spaces for Stable Variational Autoencoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks