Scaling Reinforcement Learning for Diffusion Models via Velocity Matching
cs.CV, cs.LG
Submitted: 2026-08-24
Updated: 2026-09-27
Comments: 30 pages, 13 figures
Code: https://github.com/black-forest-labs/flux
Project page: https://jaemoo-choi.github.io/RVM
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing methods largely inherit policy-gradient machinery from
Terminology
Abstract
Reward fine-tuning is becoming an important tool for adapting diffusion models to human preferences and task-specific objectives, but existing methods largely inherit policy-gradient machinery from large language models. Unlike autoregressive models, diffusion models do not provide tractable likelihoods for generated samples. As a result, current approaches either construct trajectory likelihoods from stochastic denoising transitions or approximate endpoint likelihoods with evidence lower bound, introducing additional computation and algorithmic complexity. We demonstrate that this likelihood-based machinery is not necessary for effective diffusion reward fine-tuning. We propose reward-based velocity matching (RVM), a simple trajectory-free update that acts directly on the velocity field. RVM reinforces directions associated with high-reward generations, suppresses those with low reward, and involves an optional anchor term controlling drift from a reference velocity. Notably, it provides a general framework that recovers recent fine-tuning methods, including RAM and DiffusionNFT, as special cases. Across various large-scale diffusion models reward fine-tuning tasks, RVM is competitive with or outperforms trajectory-based policy-gradient methods under substantially reduced training cost. We further find that, once the velocity update is simplified, the particular loss variant matters less than reward and anchor design. For video generation, standard preference rewards can favor visually clean but nearly static outputs; introducing a new dynamic-tracking reward that substantially improve motions while improving overall VBench performance. These results suggest that scalable reward fine-tuning for diffusion models is better posed in the native velocity representation than as likelihood-based policy optimization.
Sources
- Reinforce Adjoint Matching: Scaling RL Post-Training of Diffusion and Flow-Matching Models
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- SkyReels-V2: Infinite-length Film Generative Model
- Soft Adaptive Policy Optimization
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- AR-CoPO: Align Autoregressive Video Generation with Contrastive Policy Optimization
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in UMMs via Decompositional Verifiable Reward
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- Back to Basics: Let Denoising Generative Models Denoise
- LeapAlign: Post-Training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
- Understanding R1-Zero-Like Training: A Critical Perspective
- RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO
- The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Score-Based Generative Modeling through Stochastic Differential Equations
- V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
- Wan: Open and Advanced Large-Scale Video Generative Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models