How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning
cs.LG
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by/4.0/
The gist: Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad
Terminology
Abstract
Stopgrads are widely used in training machine learning models, but stopgrads can alter the gradient, stationary points and convergence guarantees of the original objective, which can make stopgrad training theoretically ungrounded. We introduce a stopgrad regression principle, which identifies a general template for stopgrad objectives with a closed-form characterization of stationary points and their uniqueness, unifying stopgrad objectives for flow maps, reinforcement learning, and diffusion samplers. We provide theoretical grounding for optimizing stopgrad flow map objectives by showing their unique stationary point is the true flow map, and showing positive convergence results for Eulerian and Lagrangian objectives, including MeanFlow and improved MeanFlow. Remarkably, we show that under functional semi-gradient flow, the learned flow map has a closed-form expression composing the initial flow map and the true flow map. We additionally use our stopgrad regression principle to propose modified stopgrad placements for flow map objectives which reduce training memory by 2x.
Sources
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- Flow map matching with stochastic interpolants: A mathematical framework for consistency models
- Mean Flows for One-step Generative Modeling
- Improved Mean Flows: On the Challenges of Fastforward Generative Models
- Consistency Trajectory Models: Learning Probability Flow ODE Trajectory of Diffusion
- Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
- Flow Matching for Generative Modeling
- Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models
- Approximate Temporal Difference Learning is a Gradient Descent for Reversible Policies
- Align Your Flow: Scaling Continuous-Time Flow Map Distillation
- Progressive Distillation for Fast Sampling of Diffusion Models
- Improved Techniques for Training Consistency Models
- Consistency Models
- Terminal Velocity Matching
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks