Dynamics of Gradient Descent with Large Step Size Near a Manifold of Flat Minima
cs.LG, math.DS, math.OC
Submitted: 2026-07-09
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: An important quantity in the theory of gradient descent (GD) is the sharpness, defined as the largest eigenvalue of the objective Hessian.
Terminology
Abstract
An important quantity in the theory of gradient descent (GD) is the sharpness, defined as the largest eigenvalue of the objective Hessian. Classical analyses typically require the step size to be uniformly smaller than twice the reciprocal of the sharpness, but this condition is frequently violated in the training of deep neural networks. Recent work bridges this gap in the setting of overparametrised least-squares with a single scalar output, providing a normal form for large-step GD in a neighbourhood of an isolated flat minimum and establishing three corresponding convergence results. In this paper, we extend this theory in two directions: (1) to overparametrised least-squares with vector-valued outputs (including regression with arbitrarily many observations), and (2) to a neighbourhood of a manifold of flat minima (which we show is essential for applications such as matrix factorisation). We generalise both the normal form and all three convergence theorems of to this broader setting, overcoming several technical challenges. We further show that our framework applies to deep matrix factorisation under mild assumptions, yielding several new structural results. In particular, we prove that the set of flat minima forms a fibre bundle over a product of spheres, and that the sharpness is Morse-Bott along this manifold.
Sources
- Flatness is a False Friend
- Accelerated Gradient Descent via Long Steps
- Centre manifold theorem for maps along manifolds of fixed points
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks