DreOPD: Degraded-Reference Extrapolative On-Policy Distillation for Flow-matching Models
Mingfeng Lin, Chengfei Cai, Lin Xu, Yuxiang Wei, Liang Han
Harbin Institute of Technology (Shenzhen) · Zhejiang University · Harbin Institute of Technology
cs.LG, cs.CV
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: project page: https://sleepy1231.github.io/DreOPD/
Project page: https://sleepy1231.github.io/DreOPD
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
Terminology
Summary
Summary
This paper introduces DreOPD (Degraded-reference extrapolative On-Policy Distillation), a post-training method for flow-matching models used in text-to-image generation. The method addresses the challenge of consolidating multiple task-specific strengths (e.g., prompt following, text rendering, aesthetics) into a single model that can improve beyond the individual teachers.
Core Problem and Motivation
Flow-matching models are mainstream for image generation, but adapting them to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning (RL) enables direct optimization of task-specific rewards but suffers from high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision but is imitation-based, causing a shared multi-task student to interpolate among specialized teachers rather than systematically exceed them. The paper states: RL permits extrapolation but is difficult to optimize, whereas OPD is stable but teacher-bounded.
Key Contribution: Closed-Form Velocity Target
The main technical contribution is converting trajectory-level implicit reward extrapolation into a closed-form velocity regression objective for flow-matching models. Under shared-covariance Gaussian transitions, the conditional transition objective at each student-visited state has the pointwise optimizer:
vλ⋆ = vT + (λ − 1)(vT − vref)
where vT is the teacher velocity, vref is the reference velocity, and λ ≥ 1 controls extrapolation strength. At λ = 1, this reduces to teacher imitation; for λ > 1, the target moves beyond the teacher along the teacher-reference direction. The paper notes: The teacher-reference contrast determines the direction and magnitude of extrapolation.
Degraded Reference Construction
The paper introduces a mildly degraded reference to strengthen the teacher-reference contrast. The reference is constructed from the same base model using controlled corruption such as weight quantization or mild noise injection into the velocity output. The paper states: "A reference close to the teacher yields a small contrast, making the extrapolation signal weak. This motivates using a mildly degraded reference to enlarge teacher-reference contrast while retains reward-aligned directions." Proposition 3 shows that within a reward-alignment model, a lower-alignment reference increases both the effective reward-tilt coefficient and the local rate of reward change without altering the reward direction.
Theoretical Guarantees
The paper provides two key theoretical results:
-
Lemma 1: If the teacher solves a KL-regularized reward optimization problem, then the teacher-reference log-density ratio recovers the reward up to a positive scale and additive constant.
-
Proposition 2: The expected reward under the extrapolated distribution is monotonically non-decreasing in λ, with strict inequality whenever the reward is non-constant.
Experimental Setup
Experiments use SD3.5-M at 512×512 resolution. Three task-specific teachers are used: GenEval teacher (trained with DiffusionNFT for compositional prompt following), OCR teacher (trained with GRPO-Guard for text rendering), and aesthetics teacher (trained with GRPO-Guard optimizing PickScore, ClipScore, and HPSv2.1). Baselines include Flow-OPD, DiffusionOPD, Flow-GRPO, GRPO-Guard, DiffusionNFT, and CascadeNFT.
Key Results
Single-teacher distillation: DreOPD surpasses corresponding teachers on four of five metrics:
-
GenEval: improved from 0.9470 to 0.9710
-
OCR: improved from 0.9239 to 0.9364
-
PickScore: improved from 24.034 to 24.037
-
HPSv2.1: improved from 0.3460 to 0.3497
Multi-teacher distillation: DreOPD achieves the highest average score of 0.9939, outperforming corresponding teachers on GenEval, OCR, HPSv2.1, Aesthetic, and ImageReward metrics while matching PickScore and retaining comparable ClipScore. The degraded reference raises the average score from 0.9841 to 0.9939.
Ablation Studies
-
λ factor: λ = 1.25 provides the best cross-task balance; λ = 1.5 improves OCR but degrades GenEval and perceptual quality.
-
Degraded reference: 8-bit velocity quantization achieves the highest average score; moderate degradation performs best, while overly weak degradation gives limited contrast and stronger degradation disrupts reference structure.
-
Noise level: Deterministic ODE sampling performs best, followed by noise levels of 0.3, 0.5, and 0.7; policy-gradient baseline performs worst at the same noise level.
Conclusion
The paper concludes: "DreOPD translates trajectory-level implicit reward extrapolation into a closed-form velocity target, extending OPD from teacher imitation to teacher-reference extrapolation while retaining regression-based training. We further showed that a mildly degraded reference can strengthen a reward-aligned contrast without disrupting generative structure. Across multiple settings, DreOPD achieves the best average performance over the baselines, while surpassing teachers on most metrics."
Improvements for AI systems
Improvements to AI Systems Based on DreOPD:
-
Multi-Skill Consolidation Beyond Teacher Performance: The AI system can now merge multiple specialized capabilities (e.g., prompt following, text rendering, aesthetic quality) into a single model that exceeds each individual teacher on most metrics, rather than merely interpolating between them. This is achieved by using a closed-form velocity target that extrapolates along the teacher-reference contrast direction, enabling the student to outperform all teachers simultaneously.
-
Stable Extrapolation Without High-Variance RL: The system replaces unstable policy-gradient RL with a regression-based objective, providing dense, low-variance supervision. This allows the AI to extrapolate beyond teacher rewards (e.g., improving GenEval from 0.9470 to 0.9710, OCR from 0.9239 to 0.9364) while avoiding the optimization instability typical of RL fine-tuning.
-
Controllable Extrapolation Strength via λ: The system includes a tunable hyperparameter λ that controls how far the model extrapolates beyond the teacher. This enables task-specific balancing—e.g., λ=1.25 gives the best cross-task average, while λ=1.5 can be selected for OCR-heavy applications, offering a practical knob for deployment-specific trade-offs.
-
Degraded-Reference-Driven Reward Alignment: By using a mildly degraded reference (e.g., 8-bit quantized velocities), the system amplifies the teacher-reference contrast, which strengthens the reward-aligned direction without corrupting generative structure. This improves multi-teacher distillation scores from 0.9841 to 0.9939, making the system more effective at aligning with human preferences (HPSv2.1, ImageReward) and task-specific rewards.
-
Theoretical Guarantee of Monotonic Reward Improvement: The system provides a provable guarantee that expected reward is non-decreasing with extrapolation strength λ (under non-constant rewards). This means users can confidently increase extrapolation to push performance, knowing it will not degrade the reward signal—a critical safety property for automated fine-tuning pipelines.
-
Robustness to Noise and Sampling Strategies: The system works best with deterministic ODE sampling but remains functional under moderate noise (up to 0.3), making it adaptable to stochastic inference environments. It also outperforms policy-gradient baselines at the same noise level, ensuring consistent gains even in noisy or approximate sampling settings.
-
Practical Post-Training Efficiency: The method requires only a single forward pass of the teacher and reference models to compute the closed-form target, avoiding iterative RL rollouts. This reduces computational cost and training time, enabling faster iteration on new tasks or datasets without sacrificing performance gains.
What the Improved AI System Can Do:
-
Generate images that simultaneously excel at compositional prompts, accurate text rendering, and high aesthetic appeal, surpassing any single specialized model.
-
Be fine-tuned on new tasks with stable, regression-based training that provably improves reward metrics, even when multiple objectives conflict.
-
Offer a simple, interpretable control (λ) to adjust the trade-off between conservatism (imitating teachers) and ambition (extrapolating beyond them) for different deployment scenarios.
-
Maintain high performance under noisy or approximate sampling, making it suitable for real-time or resource-constrained inference.
Abstract
Flow-matching models are now a mainstream method to image generation, but its adaptation to diverse downstream scenarios typically relies on post-training, which may cause conflicts among task-specific optimization objectives. Reinforcement learning enables direct optimization of task-specific rewards beyond the original models, yet trajectory-level optimization may incur high-variance gradients and cross-task interference. On-policy distillation (OPD) offers dense and stable supervision on student rollouts, but conventional teacher matching remains imitation-based. We propose DreOPD, a Degraded-reference extrapolative OPD method for flow-matching models that bridges these two paradigms. Our DreOPD converts implicit reward extrapolation into closed-form velocity regression, enabling extrapolative post-training with the stability of OPD. It further uses a mildly degraded reference to strengthen the teacher-reference contrast, yielding a clearer extrapolation direction. Experiments on single- and multi-teacher settings show that DreOPD outperforms OPD and multi-task RL baselines in average performance, while surpassing specialized teachers on most metrics.
Sources
- Flow-OPD: On-Policy Distillation for Flow Matching Models
- MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
- DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Aligning Text-to-Image Diffusion Models with Reward Backpropagation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- A Survey of On-Policy Distillation for Large Language Models
- Score-Based Generative Modeling through Stochastic Differential Equations
- Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models
- DanceGRPO: Unleashing GRPO on Visual Generation
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
- Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation
- MARBLE: Multi-Aspect Reward Balance for Diffusion RL
- DiffusionNFT: Online Diffusion Reinforcement with Forward Process
- DanceOPD: On-Policy Generative Field Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks