The Tamed Subgradient Unadjusted Langevin Algorithm beyond Convexity
Iosif Lytras, Nikolaos Makras, Sotirios Sabanis
cs.LG, math.OC, math.PR, stat.ML
Submitted: 2026-08-06
Comments: 53 pages
Code: https://github.com/karpathy/nanochat
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex.
Terminology
Abstract
We study the problem of sampling from target distributions whose potentials are simultaneously non-smooth, subject to superlinear gradient growth, and non-convex. We introduce the Subgradient Tamed Unadjusted Langevin Algorithm (SG-TULA), a discretisation of the Langevin diffusion that operates directly on subgradients, without relying on computationally demanding smoothing procedures. To handle the superlinear regime, taming techniques are employed to produce a stable, explicit scheme. We derive non-asymptotic convergence bounds in Wasserstein-2 distance, with all constants tracked explicitly in terms of dimension and inverse temperature, improving upon the currently known rates for subgradient-based Langevin algorithms. We further provide excess risk estimates for the associated optimisation problem. We verify the assumptions, with explicit constants, for the regularized pretraining potential of a LLM in the GPT-2 lineage and the boosted coordinate-wise variant of SG-TULA pretrains the former competitively against finetuned AdamW and Muon, for which no comparable non-asymptotic guarantees are presently available.
Sources
- The ballistic limit of the log-Sobolev constant equals the Polyak-{\L}ojasiewicz constant
- Ergodicity of Langevin Dynamics and its Discretizations for Non-smooth Potentials
- Anchored Langevin Algorithms
- DC-LA: Difference-of-Convex Langevin Algorithm
- Error estimates for tamed Euler and Randomized Euler schemes for SDEs with locally Lipschitz drift with applications to non-logconcave sampling and optimization
- kTULA: A Langevin sampling algorithm with improved KL bounds under super-linear log-gradients
- On the performance of the Euler-Maruyama scheme for multidimensional SDEs with discontinuous drift coefficient
- A fully data-driven approach to minimizing CVaR for portfolio of assets via SGLD with discontinuous updating
- When Langevin Monte Carlo Meets Randomization: New Sampling Algorithms with Non-asymptotic Error Bounds beyond Log-Concavity and Gradient Lipschitzness
- Nemotron-CLIMB: CLustering-based Iterative Data Mixture Bootstrapping for Language Model Pre-training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks