State-Dependent Lyapunov Analysis of Rank-1 Matrix Factorization
math.NA, cs.LG, cs.NA, math.DS, math.OC
Submitted: 2026-04-28
Updated: 2026-09-06
License: http://creativecommons.org/licenses/by/4.0/
The gist: We develop a state-dependent Lyapunov framework for gradient descent on rank-1 matrix factorization.
Terminology
Abstract
We develop a state-dependent Lyapunov framework for gradient descent on rank-1 matrix factorization. A parameterized quadratic certificate I(δ;, times) generates strictly nested sublevel sets. Their ordering assigns each point a state δ, while a boundary-inward property makes this state monotone along gradient-descent trajectories. Together with internal chain transitivity, this geometry identifies the limiting dynamics even when all relevant stationary points are unstable. For scalar and rank-1 factorization, it yields convergence to a global minimizer below the stability threshold and to a balanced period- 2 orbit in a post-critical interval, for almost every initialization in an explicit region. We formalize the mechanism through structural and dynamical axioms. Within the origin-centered, exchange-symmetric quadratic class, the axioms determine the ordered level-set geometry uniquely up to a monotone relabeling of the state, recovering the scalar geometry identified by Liang--Mont ú far. Additional analytic and numerical examples suggest broader applicability.
Sources
- Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
- Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
- The Zero Set of a Real Analytic Function
- Understanding Edge-of-Stability Training Dynamics with a Minimalist Example
Related papers
- Do physics-informed neural networks (PINNs) need to be deep? Shallow PINNs using the Levenberg-Marquardt algorithm
- A Neural-preconditioned Poisson Solver for Mixed Dirichlet and Neumann Boundary Conditions
- Second-order consistency for learning chaotic dynamics via randomized Jacobian matching
- Windowed thinning and query complexity for the bouncy particle and Zigzag samplers
- Data-efficient Kernel Methods for Learning Hamiltonian Systems
- Adjoint Method versus Physics-Informed Neural Networks in PDE-Constrained Inverse Problems