Stable Velocity: A Variance Perspective on Flow Matching
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Stable Velocity: A Variance Perspective on Flow Matching".
Tom: Stable Velocity introduces a variance-based perspective on flow matching, revealing a two-regime structure that governs both training and inference dynamics.
Jane: First, who's behind it and why it matters.
Paper summary: Jane: Okay, so as we move into the specifics of "Stable Velocity," the authors introduce three main ideas to handle those different variance levels. They call these Stable Velocity Matching for training, Variance-Aware Representation Alignment for selective supervision, and Stable Velocity Sampling for inference acceleration. It seems like a unified approach is being proposed here.
Tom: That's right; the thesis centers on proposing this unified framework to solve the high variance problem inherent in single-sample conditional velocities. They claim that Stable Velocity Matching creates an unbiased variance-reduction objective, while VA-REPA smartly applies auxiliary supervision only when the variance is low, and StableVS uses that same low-variance structure for faster sampling.
Lu: The authors are proposing replacing a single target with a multi-sample, self-normalized aggregation over reference data points under the multi-sample conditional path for StableVM; it’s a mathematical shift in how we define the training objective Lu. This is essentially how they propose to keep the global minimizer the same while drastically cutting down on training variance.
Meng: From an engineering standpoint, replacing one sample target with an aggregation sounds computationally intensive; how do they manage that complexity without slowing down our hardware significantly during a standard run?
Lalam: Lalam looks at this and sees that because StableVM keeps the global minimizer the same as CFM, it means we get the benefits of variance reduction without having to completely overhaul our existing, well-tested training pipelines. This is very pragmatic for deployment.
Tom: Precisely; they show Theorem three point one(a) proving that this StableVM target remains unbiased and shares the same global minimizer as CFM Tom. And then, in terms of inference, StableVS exploits the fact that in the low-variance regime, the instantaneous velocity is mostly determined by a single dominant data point.
Jane: That's where they show a closed-form solution for linear interpolants using Stable Velocity Sampling to enable finetuning-free acceleration Jane. It’s not just about reducing noise; it’s about finding shortcuts in the math when we are in that specific low-variance zone.
Lu: And the empirical validation confirms this, showing that StableVM and VA-REPA consistently outperform prior REPA methods like REPA-E across different model scales Lu. The authors also highlight how the split point xi, set to zero point seven as a default, balances performance by preventing noisy supervision from the high-variance regime from dragging things down Lu.
Meng: So, for practical application, it sounds like we get better training stability and faster sampling without needing to retrain everything from scratch or spend massive amounts of time tuning hyperparameters manually.
Lalam: Lalam agrees; if this framework can deliver more stable training with less effort on our side, it means we can deploy these flow-based models more quickly into real-world applications.
Conclusion: Tom: So, wrapping up this discussion on "Stable Velocity: A Variance Perspective on Flow Matching," the core message from Donglin Yang et al. is that explicitly modeling the variance structure along the generative trajectory gives us a principled way to design training objectives and sampling algorithms Tom. They’ve shown that by splitting the process into high-variance and low-variance regimes, we can tailor our methods specifically for each one.
Jane: What this means simply is that instead of applying one blanket training technique, you use a specialized objective in the noisy regions and a different, much more efficient method when you're close to the real data distribution Jane. The authors are linking these concepts—StableVM, VA-REPA, and StableVS—under one coherent principle of variance control.
Lu: The implication for future research is that we now have a clear direction: instead of just aiming for a globally optimized loss, we should be thinking about how to manage the variance landscape across the entire generative path Lu. This opens up avenues for developing new types of regularization techniques based on this regime detection.
Meng: From an engineering perspective, I see it suggesting that we can build hybrid systems where different parts of our pipeline dynamically switch between these methods based on where the generation is happening, which could lead to more robust applications.
Lalam: Lalam thinks the impact here is really about making AI systems fundamentally more predictable in their behavior; when we understand the variance structure, we gain control over instability during training and inference Lalam. This kind of stability translates directly into user trust and reliability for any application powered by these models.
Tom: Exactly, so it’s not just a technical tweak; it’s a structural understanding of the generative process that allows us to build more robust AI components Tom. The authors successfully showed that this variance perspective leads to consistent improvements in training stability and sampling speedups without compromising sample quality across various benchmarks.
Jane: It seems the real value of this work lies in moving flow matching from a black-box optimization problem toward a structured, controllable system where we can anticipate and manage how the AI behaves during both learning and generation Jane. This provides a solid foundation for next-generation generative models.
Donglin Yang, Yongxing Zhang, Xin Yu, Liang Hou, Xin Tao, Pengfei Wan, Xiaojuan Qi
University of Hong Kong · University of British Columbia
cs.CV
Submitted: 2026-02-05
Updated: 2026-09-28
Code: https://github.com/linYDTHU/StableVelocity
Importance score: 92/100
The gist: Stable Velocity introduces a variance-based perspective on flow matching, revealing a two-regime structure that governs both training and inference dynamics.
Key concepts
- Two-Regime Structure
- Flow matching targets exhibit two distinct variance regimes: one where the data distribution is far from the prior (high variance) and another where it is close to the data distribution (low variance). This structural difference dictates how training noise behaves and informs targeted optimization strategies.
- Stable Velocity Matching (StableVM)
- This objective replaces single-sample velocity targets with a multi-sample, self-normalized aggregation of velocities. It remains unbiased while strictly reducing training variance compared to standard flow matching, ensuring the global minimizer is preserved.
- Variance-Aware Representation Alignment (VA-REPA)
- This technique selectively applies auxiliary supervision based on the variance regime. It uses a weighting function to scale alignment losses, ensuring supervision is effective and prevents vanishing gradients when samples fall into the high-variance region.
Terminology
Summary
Stable Velocity introduces a variance-based perspective on flow matching, revealing a two-regime structure that governs both training and inference dynamics. This framework addresses the high variance in single-sample conditional velocities by proposing Stable Velocity Matching (StableVM) for unbiased variance reduction during training, Variance-Aware Representation Alignment (VA-REPA) for selective auxiliary supervision, and Stable Velocity Sampling (StableVS) for finetuning-free acceleration at inference.
Variance Analysis of Flow Matching
The paper first establishes the mathematical foundation by reviewing flow matching and stochastic interpolants. It defines the continuous-time corruption process as a Stochastic Differential Equation (SDE) where the conditional velocity field is derived from this process. The key insight is characterizing the variance of these targets, quantified by VCFM(t), which exhibits two distinct regimes: a low-variance regime near the data distribution
and a high-variance regime near the prior.
This structure naturally suggests two primary research questions: (1) How to reduce training variance in the high-variance regime without altering the global minimizer, and (2) How to exploit the low-variance regime for stronger supervision and faster sampling.
Variance-Driven Optimization of Training and Sampling
The framework introduces three core components derived from this analysis:
-
Stable Velocity Matching (StableVM): This is an
unbiased variance-reduction objective
that replaces the single-sample conditional velocity target with amulti-sample, self-normalized aggregation over reference data points under the multi-sample conditional path.
The StableVM target, denoted as vbStableVM, is defined as the self-normalized importance weighted average of conditional velocities. A key result is Theorem 3.1(a), which proves that the StableVM targetremains unbiased
and admitsthe same global minimizer as CFM.
Furthermore, it shows that VStableVM strictly reduces variance compared to CFM, with a bound showing it is bounded by the variance of CFM plus a term related to the difference between true and conditional velocities. -
Variance-Aware Representation Alignment (VA-REPA): This component adapts auxiliary supervision selectively based on the variance regime. The paper empirically finds that
the effectiveness of REPA largely arises when applied in the low-variance regime,
as semantic alignment iswell-conditioned
there, whereas it saturates in the high-variance regime. VA-REPA introduces a non-negative weighting function w(t) to modulate the representation alignment loss, ensuring supervision is scaled by the number of effective samples, preventing vanishing gradients when most samples fall in the high-variance regime. -
Stable Velocity Sampling (StableVS): This provides a
finetuning-free acceleration
strategy for inference specifically targeting the low-variance regime. In this regime, becausethe instantaneous velocity vt(xt) is effectively determined by a single dominant data point x0,
StableVS exploits this structure to enablestable, large-step integration without degrading sample quality.
For the linear interpolant case, this leads to a closed-form solution:xτ = xt + (τ − t)vt(xt).
Experimental Validation
The proposed framework is validated across various models and benchmarks. StableVM and VA-REPA consistently outperform prior REPA methods (like REPA-E) in terms of FID, IS, precision, and recall across different model scales. The ablation studies indicate that the split point ξ (set to 0.7 as a default) balances performance: a smaller split point yields better results at early training stages, while a larger one degrades performance due to noisy supervision from the high-variance regime. StableVS demonstrates significant acceleration, achieving more than 2× inference acceleration in the low-variance regime
for models like SD3.5 and Flux without perceptible degradation in sample quality.
Conclusion
The work concludes that explicitly modeling variance structure along the generative trajectory provides a principled foundation for designing more efficient training objectives and sampling algorithms. The Stable Velocity framework—comprising StableVM, VA-REPA, and StableVS—unifies variance reduction and auxiliary supervision under a single principle, leading to consistent improvements in training stability and substantial sampling speedups without sacrificing sample quality.
The gist
Stable Velocity reveals a two-regime structure governing flow matching dynamics: a high-variance regime near the prior where optimization is noisy, and a low-variance regime near the data distribution where conditional and marginal velocities coincide. StableVM reduces training variance through unbiased aggregation over reference samples, VA-REPA applies auxiliary supervision selectively in the low-variance regime, and StableVS enables finetuning-free acceleration by exploiting deterministic dynamics in that same low-variance region.
How it works
-
StableVM replaces the single-sample conditional velocity target with a
multi-sample, self-normalized aggregation over reference data points under the multi-sample conditional path
to create an unbiased, variance-reduced objective.
Improvements for AI systems
Based on the provided scientific paper, here are specific, actionable improvements for existing AI systems and what those improved systems can achieve:
) 1. Implement Stable Velocity Matching (StableVM) for Training:
-
Replace standard Conditional Flow Matching (CFM) objectives with the StableVM objective in training. This involves using a multi-sample, self-normalized aggregation of conditional velocities from a reference batch instead of single-sample estimates.
-
Use Variance-Aware Representation Alignment (VA-REPA) adaptively during training by applying it selectively in the low-variance regime (e.g., at early timesteps).
"Improved System Capability: The system will exhibit significantly improved optimization stability and convergence speed, especially when training on high-resolution data or complex conditional tasks. By reducing the variance of the training targets, the model will train more reliably without requiring extensive hyperparameter tuning for noise sensitivity."
) 2. Enable Finetuning-Free Inference Acceleration via Stable Velocity Sampling (StableVS):
- During inference in the low-variance regime (identified by a specific timestep range, e.g., [0, 0.7]), replace standard numerical solvers (like Euler or DPM++) with the StableVS strategy. This involves using closed-form simplifications derived from the deterministic dynamics of the flow in this regime to enable large, finetuning-free integration steps (e.g., 9 steps instead of 30).
"Improved System Capability: The system will achieve over 2x faster sampling speeds during inference for generative tasks (images and videos) without sacrificing sample quality. This allows for real-time or near real-time generation, drastically reducing the computational latency associated with high-step diffusion models."
) 3. Optimize Training Dynamics via Variance Regime Awareness:
- Dynamically determine the variance regime based on the current timestep and data dimensionality (e.g., using a learned or empirically determined split point ξ). Apply StableVM for training in the high-variance regime ([ξ, 1]) and activate VA-REPA only in the low-variance regime ([0, ξ]).
"Improved System Capability: The system will exhibit superior performance across diverse data scales (e.g., ImageNet 256x256) and modalities (text-to-image/video). By tailoring the supervision strategy to where it is most informative, the model will achieve better FID and IS scores than methods that apply uniform supervision, leading to more robust and high-quality outputs."
) 4. Enhance Conditional Generation for Complex Prompts:
- Extend StableVM's class-conditional memory bank mechanism (Algorithm 2) to handle classifier-free guidance (CFG). This allows the model to construct the mixture input and target field using a diverse, class-specific reference set even when per-batch class frequencies are low.
"Improved System Capability: The system will excel at generating images and videos based on complex, multi-faceted prompts (e.g., 'A dog plays guitar while a cat takes a selfie'). This capability is enhanced by the ability to robustly handle conditional constraints through the memory bank, leading to higher precision and recall in complex generation tasks."
) 5. Enable Deterministic Sampling for Flow Models:
- For flow-based models, leverage StableVS's closed-form solutions when sampling in the low-variance regime. This allows for exact integration via Euler steps of arbitrary size (e.g., xτ = xt + (τ - t)vt(xt)).
"Improved System Capability: The system will provide highly accurate and predictable trajectory following during the early stages of generation, ensuring that the generated content adheres precisely to the learned flow dynamics in this regime."
Sources
- Advances in Importance Sampling
- Classifier-Free Diffusion Guidance
- REPA-E: Unlocking VAE for End-to-End Tuning with Latent Diffusion Transformers
- Flow Matching for Generative Modeling
- Flow Matching Guide and Code
- Progressive Distillation for Fast Sampling of Diffusion Models
- What matters for Representation Alignment: Global Information or Spatial Structure?
- Score-Based Generative Modeling through Stochastic Differential Equations
- Improving and generalizing flow-based generative models with minibatch optimal transport
- Wan: Open and Advanced Large-Scale Video Generative Models
- Diffuse and Disperse: Image Generation with Representation Regularization
- REPA Works Until It Doesn't: Early-Stopped, Holistic Alignment Supercharges Diffusion Training
- Qwen-Image Technical Report
- Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- Fast Training of Diffusion Models with Masked Transformers
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models