Muon meets Tamed Langevin: Momentum Preconditioning beyond Convex and gradient-Lipschitz Potentials
cs.LG, math.OC, math.PR, stat.ML
Submitted: 2026-10-01
Updated: 2026-10-01
Code: https://github.com/karpathy/nanochat
Terminology
Sources
- On MUON optimization: From non-convergence to an error analysis with Polar Express and the Newton-Schulz polynomial from implementations
- Trend to equilibrium and Newtonian limit for the relativistic Langevin equation with singular potentials
- Softsign: Smooth Sign in Your Optimizer For Better Parameter Heterogeneity Handling
- Kimi K2: Open Agentic Intelligence
- Muon is Scalable for LLM Training
- Musec: MomentUm SpEctral Clipping for Stable Muon-type Training
- The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
- Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer
- Muon Does Not Converge on Convex Lipschitz Functions
- Training Deep Learning Models with Norm-Constrained LMOs
- Delving into Muon and Beyond: Deep Analysis and Extensions
- Practical Efficiency of Muon for Pretraining
- Spectral Scaling Laws of Muon
- HTMuon: Improving Muon via Heavy-Tailed Spectral Correction
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks