Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
cs.LG
Submitted: 2025-09-29
Updated: 2026-09-01
Comments: Accepted at ICML 2026; updated to the camera-ready version
Code: https://github.com/scxue/advantage
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Stochastic Interpolants: A Unifying Framework for Flows and Diffusions
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Training Diffusion Models with Reinforcement Learning
- Directly Fine-Tuning Diffusion Models on Differentiable Rewards
- Adjoint Matching: Fine-tuning Flow and Diffusion Generative Models with Memoryless Stochastic Optimal Control
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- TempFlow-GRPO: When Timing Matters for GRPO in Flow Models
- Classifier-Free Diffusion Guidance
- Aligning Text-to-Image Models using Human Feedback
- Reward Guided Latent Consistency Distillation
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- BranchGRPO: Stable and Efficient GRPO with Structured Branching in Diffusion Models
- Flow Matching for Generative Modeling
- Flow-GRPO: Training Flow Matching Models via Online RL
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
- Flow Matching Policy Gradients
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- Non-Denoising Forward-Time Diffusions
- Hierarchical Text-Conditional Image Generation with CLIP Latents
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks