A Theory of Saddle Escape in Deep Nonlinear Networks
cs.LG, cond-mat.dis-nn, stat.ML
Submitted: 2026-05-02
Updated: 2026-09-25
Terminology
Sources
- SGD learning on neural networks: leap complexity and saddle-to-saddle dynamics
- The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
- High-dimensional dynamics of generalization error in neural networks
- Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
- On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization
- Implicit Regularization in Deep Matrix Factorization
- High-dimensional limit theorems for SGD: Effective dynamics and critical scaling
- Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
- Weight-space symmetry in deep networks gives rise to permutation saddles, connected by equal-loss valleys across the loss landscape
- On Lazy Training in Differentiable Programming
- Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity
- Gradient descent aligns the layers of deep linear networks
- How to Escape Saddle Points Efficiently
- Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics
- Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
- Saddle-to-Saddle Dynamics in Diagonal Linear Networks
- On the Spectral Bias of Neural Networks
- Exact solutions to the nonlinear dynamics of learning in deep linear neural networks
- The Neural Race Reduction: Dynamics of Abstraction in Gated Networks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks