Gradient Descent with Stochastic Subspaces via Persistence of Memory
math.OC, cs.LG, eess.SP, stat.ML
Submitted: 2026-09-16
Updated: 2026-09-16
Comments: 81 pages, 2 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Randomised subspace methods for non-convex optimization, with applications to nonlinear least-squares
- A Memory Efficient Randomized Subspace Optimization Method for Training Large Language Models
- Gradient Descent Happens in a Tiny Subspace
- Randomized Subspace Nesterov Accelerated Gradient
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Federated Full-Parameter Tuning of Billion-Sized Language Models with Communication Cost under 18 Kilobytes
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Ferret: Federated Full-Parameter Tuning at Scale for Large Language Models
- Randomized Forward Mode of Automatic Differentiation For Optimization Algorithms
- GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection
Related papers
- Lions and Muons: Optimization via Stochastic Frank-Wolfe under Heavy-Tailed Noise
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- Incremental Learning in Mirror Flows
- Online Control via Counterfactual Tracking
- Asynchronous Replanning in Two Population Linear Quadratic Mean Field Games: Information Requirements and Stability
- Petrov-Galerkin operator inference with application to stability-encouraging identification