A Flatness-Generalization Relation in the Teacher-Student Tree-Committee Machine
cond-mat.dis-nn, cs.LG, stat.ML
Submitted: 2026-09-25
Updated: 2026-09-25
Terminology
Sources
- On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima
- Why flatness does and does not correlate with generalization for deep neural networks
- Flatness After All?
- A Modern Look at the Relationship between Sharpness and Generalization
- Relative Flatness and Generalization
- Flatness and Generalization: Learning Multi-Index Models with Homogeneous Neural Networks
- Empirical Analysis of the Hessian of Over-Parametrized Neural Networks
- Unveiling the Hessian's Connection to the Decision Boundary
- Emergent properties of the local geometry of neural loss landscapes
- Local minima of the empirical risk in high dimension: General theorems and convex examples
- Topological trivialization in non-convex empirical risk minimization
- The Role of the Time-Dependent Hessian in High-Dimensional Optimization
- Overparametrization bends the landscape: BBP transitions at initialization in simple Neural Networks
- Generalization performance of narrow one-hidden layer networks in the teacher-student setting
Related papers
- Few-Shot Neuromorphic Vision in a Nonlinear Photonic Network Laser
- Hyperbolic lattices with mass disorder: Phases and phase transitions
- Adaptive Neural Quantum States: A Recurrent Neural Network Perspective
- Machine learning Majorana topology using unsupervised and supervised learning
- Quenched fluctuation-induced force arising from polarization disorder