PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Findings of EMNLP 2026; Code is available at https://github.com/VennTum99/PLC-DPO
Code: https://github.com/VennTum99/PLC-DPO
License: http://creativecommons.org/licenses/by/4.0/
The gist: Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable.
Terminology
Abstract
Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.
Sources
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Llama 3 Herd of Models
- ULMA: Unified Language Model Alignment with Human Demonstration and Point-wise Preference
- Mistral 7B
- Qwen2.5 Technical Report
- Decoupled Weight Decay Regularization
- Proximal Policy Optimization Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks