Estimating Population-Risk Curves Along Nonconvex Gradient Flows from the Training Sample

arXiv:2608.30261 · stat.ML, cs.LG · Submitted 2026-08-31 · Read on arXiv

stat.ML, cs.LG

Submitted: 2026-08-31

Updated: 2026-08-31

Comments: 63 pages

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

The gist: We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample.

Terminology

Abstract

We estimate the conditional population-risk curve of a realized smooth nonconvex gradient flow from the training sample. Flow approximate leave-one-out (Flow-ALO) propagates a deletion response and evaluates omitted observations at approximate deleted paths. The risk-curve error decomposes into response approximation, exact-LOO fluctuation, and deletion-to-full risk transfer. On each fixed finite horizon, bounded centered training-loss gradients, a one-sided Hessian lower bound, locally Lipschitz Hessians, and a strict tube-closure condition yield an explicit (n-1)-2 bound for the deletion-response error. Bounded evaluation-loss gradients transfer the deletion-response bound to the score without requiring the Hessian to be invertible. Direct first-order jackknife cancellation and exact-LOO concentration control deletion-to-full risk transfer and fluctuation, respectively, completing recovery of the conditional population-risk curve. For bounded smooth two-layer mean-field networks training both layers, the score-error bound is uniform in width.

Sources

Related papers