DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
cs.LG, cs.CV
Submitted: 2026-05-14
Updated: 2026-09-11
Project page: https://quanhaol.github.io/DiffusionOPD-site
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization.
Terminology
Abstract
Reinforcement learning has emerged as a powerful tool for improving diffusion-based text-to-image models, but existing methods are largely limited to single-task optimization. Extending RL to multiple tasks is challenging: joint optimization suffers from cross-task interference and imbalance, while cascade RL is cumbersome and prone to catastrophic forgetting. We propose DiffusionOPD, a new multi-task training paradigm for diffusion models based on Online Policy Distillation (OPD). DiffusionOPD first trains task-specific teachers independently, then distills their capabilities into a unified student along the student own rollout trajectories. This decouples single-task exploration from multi-task integration and avoids the optimization burden of solving all tasks jointly from scratch. Theoretically, we lift the OPD framework from discrete tokens to continuous-state Markov processes, deriving a closed-form per-step KL objective that unifies both stochastic SDE and deterministic ODE refinement via mean-matching. We formally and empirically demonstrate that this analytic gradient provides lower variance and better generality compared to conventional PPO-style policy gradients. Extensive experiments show that DiffusionOPD consistently surpasses both multi-reward RL and cascade RL baselines in training efficiency and final performance, while achieving state-of-the-art results on all evaluated benchmarks.
Sources
- Process Reinforcement through Implicit Rewards
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Consistency Models
- V-GRPO: Online Reinforcement Learning for Denoising Generative Models Is Easier than You Think
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- Unified Reward Model for Multimodal Understanding and Generation
- GDRO: Group-level Reward Post-training Suitable for Diffusion Models
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation
- RewardDance: Reward Scaling in Visual Generation
- UFOGen: You Forward Once Large Scale Text-to-Image Generation via Diffusion GANs
- DanceGRPO: Unleashing GRPO on Visual Generation
- One-step Diffusion with Distribution Matching Distillation
- DiffusionNFT: Online Diffusion Reinforcement with Forward Process
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks