Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
cs.LG
Submitted: 2025-10-15
Updated: 2026-09-26
Code: https://github.com/ClaraBing/Adam_or_GN
Terminology
Sources
- The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
- Towards Quantifying the Preconditioning Effect of Adam
- Pareto Frontiers in Neural Feature Learning: Data, Compute, Width, and Luck
- Adam: A Method for Stochastic Optimization
- Understanding Adam Requires Better Rotation Dependent Assumptions
- Normalized Gradients for All
- Toward Understanding Why Adam Converges Faster Than SGD for Transformers
- Per-example gradients: a new frontier for understanding and improving optimizers
- Eigenvalues of the Hessian in Deep Learning: Singularity and Beyond
- SOAP: Improving and Stabilizing Shampoo using Adam
- Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
- PyHessian: Neural Networks Through the Lens of the Hessian
- Deconstructing What Makes a Good Optimizer for Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks