Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample
stat.ML, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 63 pages
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample.
Terminology
Abstract
We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO fluctuation, and deletion-to-full risk transfer. On each fixed finite horizon, bounded centered training-loss gradients, a one-sided Hessian lower bound, locally Lipschitz Hessians, and a strict tube-closure condition yield an explicit (n-1)-2 bound for the deletion-response error. Bounded evaluation-loss gradients transfer the deletion-response bound to the score without requiring the Hessian to be invertible. Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve. For bounded smooth two-layer mean-field networks training both layers, the score-error bound is uniform in width.
Sources
- From Cross-Validation to SURE: Asymptotic Risk of Tuned Regularized Estimators
- Concentration inequalities for leave-one-out cross validation
- Optimal Confidence Band for Kernel Gradient Flow Estimator
- A Higher-Order Swiss Army Infinitesimal Jackknife
- Precise gradient descent training dynamics for finite-width multi-layer neural networks
- Statistical Inference on Gradient Flows
- A Theory of Generalization in Deep Learning
- A Dual Representation of Influence Functions for Linearizable Models
- Gradient-Flow Optimization as Dynamic Random-Effects Inference: Testing and Early Stopping with Applications to Deep Learning
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey