Eigenvalues of the Hessian in Deep Learning: The Origin of Symmetry and Its Breaking
cs.LG, math.OC, stat.ML
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- Hidden Minima in Two-Layer ReLU Networks
- Symmetry & Critical Points
- On the Principle of Least Symmetry Breaking in Shallow ReLU Models
- Symmetry Breaking in Symmetric Tensor Decomposition
- Relational inductive biases, deep learning, and graph networks
- A Teacher-Student Perspective on the Dynamics of Learning Near the Optimal Point
- Emergent properties of the local geometry of neural loss landscapes
- On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs
- Beyond Random Matrix Theory for Deep Networks
- Gradient Descent Happens in a Tiny Subspace
- Sharp description of local minima in the loss landscape of high-dimensional two-layer ReLU neural networks
- Demystifying Spectral Bias on Real-World Data
- Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
- Symmetry-via-Duality: Invariant Neural Network Densities from Parameter-Space Correlators
- Generalization Guarantees for Neural Networks via Harnessing the Low-rank Structure of the Jacobian
- The Full Spectrum of Deepnet Hessians at Scale: Dynamics with SGD Training and Sample Size
- Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
- The general theory of permutation equivarant neural networks and higher order graph variational encoders
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks