Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design
cs.LG, cs.AI
Submitted: 2026-02-04
Updated: 2026-09-17
Comments: 25 pages, 11 figures
Code: https://github.com/black-forest-labs/flux
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning has been widely applied to diffusion and flow models for visual tasks such as text-to-image generation.
Terminology
Abstract
Reinforcement learning has been widely applied to diffusion and flow models for visual tasks such as text-to-image generation. However, these tasks remain challenging because diffusion models have intractable likelihoods, which creates a barrier for directly applying popular policy-gradient type methods. Existing approaches primarily focus on crafting new objectives built on already heavily engineered LLM objectives, using ad hoc estimators for likelihood, without a thorough investigation into how such estimation affects overall algorithmic performance. In this work, we provide a systematic analysis of the RL design space by disentangling three factors: i) policy-gradient objectives, ii) likelihood estimators, and iii) rollout sampling schemes. We show that adopting an evidence lower bound (ELBO) based model likelihood estimator, computed only from the final generated sample, is the dominant factor enabling effective, efficient, and stable RL optimization, outweighing the impact of the specific policy-gradient loss functional. We validate our findings across multiple reward benchmarks using SD 3.5 Medium, and observe consistent trends across all tasks. Our method improves the GenEval score from 0.24 to 0.95 in 90 GPU hours, which is 4.6 times more efficient than FlowGRPO and 2 times more efficient than the SOTA method without reward hacking.
Sources
- GPT-4 Technical Report
- Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
- Training Diffusion Models with Reinforcement Learning
- Classifier-Free Diffusion Guidance
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- NFT: Bridging Supervised Learning and Reinforcement Learning in Math Reasoning
- Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
- The Art of Scaling Reinforcement Learning Compute for LLMs
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Soft Adaptive Policy Optimization
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Back to Basics: Let Denoising Generative Models Denoise
- Flow Matching for Generative Modeling
- Flow-GRPO: Training Flow Matching Models via Online RL
- Proximal Policy Optimization Algorithms
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Understanding R1-Zero-Like Training: A Critical Perspective
- Demystifying Diffusion Objectives: Reweighted Losses are Better Variational Bounds
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks