Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference Optimization
cs.CV, cs.AI
Submitted: 2026-04-27
Updated: 2026-09-30
Comments: 21 pages. Published as a conference paper at ICLR 2026
Code: https://github.com/L-CodingSpace/semi-dpo
Project page: https://liming-ai.github.io/SemiDPO
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment.
Terminology
Abstract
Human visual preferences are inherently multi-dimensional, encompassing aesthetics, detail fidelity, and semantic alignment. However, existing datasets provide only single, holistic annotations, resulting in severe label noise: images that excel in some dimensions but are deficient in others are simply marked as winner or loser. We theoretically demonstrate that compressing multi-dimensional preferences into binary labels generates conflicting gradient signals that misguide Diffusion Direct Preference Optimization (DPO). To address this, we propose Semi-DPO, a semi-supervised approach that treats consistent pairs as clean labeled data and conflicting ones as noisy unlabeled data. Our method starts by training on a consensus-filtered clean subset, then uses this model as an implicit classifier to generate pseudo-labels for the noisy set for iterative refinement. Experimental results demonstrate that Semi-DPO achieves state-of-the-art performance and significantly improves alignment with complex human preferences, without requiring additional human annotation or explicit reward models during training. We will release our code and models at: https://github.com/L-CodingSpace/semi-dpo
Sources
- Training Diffusion Models with Reinforcement Learning
- Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Margin-aware Preference Optimization for Aligning Diffusion Models without Reference
- DivideMix: Learning with Noisy Labels as Semi-supervised Learning
- Wuerstchen: An Efficient Architecture for Large-Scale Text-to-Image Diffusion Models
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Score-Based Generative Modeling through Stochastic Differential Equations
- Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis
- Human Preference Score: Better Aligning Text-to-Image Models with Human Preference
- Understanding deep learning requires rethinking generalization
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models