Preference Optimization with Multi-Sample Comparisons
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Preference Optimization with Multi-Sample Comparisons".
Tom: The gist:
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So the paper explains how these multi-sample comparisons work mathematically, defining the likelihood of one group being preferred over another using a sigmoid function based on the difference between their rewards.
Jane: That formula, p(Gw≽ Gl x) = Φ(r(Gw, x) − r(Gl, x)), shows how they are quantifying that preference across multiple responses at once.
Lu: And then they show the objective functions for both mDPO and mIPO, LmDPO and LmIPO, which are built to maximize this group-wise likelihood under certain constraints.
Meng: The math is complex, but it’s important because it sets up the optimization problem so that the model learns to favor distributions that have more desirable collective traits.
Tom: The key improvement they highlight is exactly what we discussed: these methods are better at optimizing diversity and bias than their single-sample predecessors.
Jane: They show empirical results across different tasks, like random number generation where mIPO and mDPO both beat the SFT baseline in uniformity.
The paper's summary: Tom: The paper shows that multi-sample comparison is genuinely more effective at optimizing collective characteristics than just a single sample comparison across various generative models.
Lu: They tested this in random number generation, and both mIPO and mDPO achieved higher uniformity compared to the SFT baseline, with mDPO showing a win rate of zero point nine nine against it <ref:2410.12138#pg1>.
Jane: That’s really telling for randomness; it means the model is much better at producing outputs that look more uniformly distributed across its possibilities than before.
Meng: For text-to-image generation, they demonstrated that multi-sample optimization significantly reduces biases related to gender and race by comparing groups of samples.
Tom: Specifically, mDPO improved the Simpson Diversity Index over the original SD one point five model for both gender and race metrics, showing a significantly higher diversity score in those areas.
Lu: In creative fiction generation, they found that mIPO with k equals five achieved a win rate of zero point three five three against DPO, which suggests improved quality and diversity at both lexical and semantic levels.
Jane: So these results suggest that when you look at the whole distribution of outputs, you get much better control over things like fairness in images or variety in stories.
The paper's improvements: Tom: The main conclusion is that multi-sample comparison is much more advantageous than single-sample comparison when dealing with distributional properties like diversity and bias.
Lu: It confirms that judging a model by looking at the whole distribution of its outputs gives us a richer understanding than just picking the single best answer out of ten.
Meng: For practical application, this suggests that if we want models that are truly robust across many scenarios, focusing on these collective characteristics is the right direction.
Jane: The paper also pointed out some technical details about stochastic estimators for mIPO, showing how you can compute unbiased estimates of the squared difference between expectations.
Tom: And they noted that as the sample size k increases, the error in their estimator actually decreases and gets closer to being unbiased, which is a good stability factor.
Lu: But there's a caveat they mentioned too: the analysis of that variance shows that even with increasing k, the error doesn't vanish instantly.
Meng: They also highlighted how this approach is especially robust against label noise, with the gap between win and lose rates narrowing as k increases when comparing default labeling versus GPT-4o labeling <ref:2410.12138#pg2>.
Jane: So we end up saying that multi-sample DPO and mIPO are powerful extensions for improving collective characteristics, provided you use enough samples to stabilize the estimation.
Tom: That’s our look at "Preference Optimization with Multi-Sample Comparisons." It shows us that looking at groups of samples is a way to get a more complete picture of what an AI model can actually produce.
Conclusion: Tom: So we’ve been looking at "Preference Optimization with Multi-Sample Comparisons," and the big takeaway is that using multiple samples to compare preferences gives us a much better idea of what a generative model is really doing collectively.
Jane: Exactly, Tom. It means instead of just judging one sentence or one image, we can look at the whole family of outputs and see things like diversity and bias much more accurately across different tasks.
Tom: Right. We saw some cool numbers where mDPO and mIPO outperformed the SFT baseline in random number generation uniformity, getting that distribution way closer to what you’d expect from a truly random source.
Lu: I think the real excitement here is how it handles those distributional problems—it’s not just about making one output better, it's about shaping the entire output landscape in a more controlled way.
Meng: From an engineering standpoint, that robustness against label noise is pretty practical; if our training data has some messy labels, these methods seem like they hold up better when you use multi-sample comparisons.
Lalam: I’m seeing how this applies to cultural generation. If we can get the model to generate more diverse narratives or images without it just defaulting to the most common patterns, that opens up so much new creative space for users.
Jane: It seems like a lot of this is going to help us build AI systems that aren't just smart, but genuinely more varied and less biased in how they express themselves.
Tom: So we’ve covered mDPO and mIPO, showing how comparing groups rather than singles helps optimize those collective traits for everything from image generation to fiction writing.
Lu: It really puts the focus on the whole model behavior, not just a single point of failure in its training.
Meng: I'm curious how we can scale this up when we have massive datasets; it seems like the overhead of comparing groups gets bigger as k goes up, so efficiency is still important.
Lalam: But if we think about culture and how AI interacts with people, having a model that shows more genuine diversity in its creative outputs really matters for building trust.
Jane: It’s a solid paper showing how to move beyond single-sample alignment toward optimizing the actual performance of the entire model distribution.
Chaoqi Wang, Zhuokai Zhao, Chen Zhu, Karthik Abinav Sankararaman, Michal Valko, Sara Cao, Zhaorun Chen, Madian Khabsa, Yuxin Chen
Meta GenAI
cs.LG, cs.CL, stat.ML
Submitted: 2024-10-16
Updated: 2025-03-26
Comments: Code is available at https://github.com/alecwangcq/multi-sample-alignment
Code: https://github.com/alecwangcq/multi-sample-alignment
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 82/100
The gist: The gist: Multi-sample Direct Preference Optimization (mDPO) and Multi-sample Identity Preference Optimization (mIPO) extend traditional post-training alignment methods by utilizing multi-sample
Key concepts
- Single-sample Comparison Limitations
- Traditional preference methods only compare two individual outputs at a time. This limits their ability to assess broader qualities of a model's output, such as overall diversity or systemic biases. They fail to capture how a model performs across many different examples simultaneously.
- Multi-sample Comparison
- mDPO and mIPO use multiple samples in their comparisons to evaluate the likelihood that one group (Gw) is preferred over another (Gl). This allows the optimization process to focus on capturing distributional characteristics, like gender or race bias, across a wider range of outputs.
- Generative Diversity and Bias
- These are critical model properties that measure how varied and fair a model's output is. Single-sample methods often miss these issues. Multi-sample optimization proves more effective at improving diversity in creative tasks and reducing harmful biases in image generation.
- Objective Functions (LmDPO/LmIPO)
- These are mathematical goals used to train the model. For mDPO, it maximizes a reward function under certain constraints, while for mIPO, it focuses on minimizing the difference between expected outcomes of preferred and non-preferred groups across many samples.
Terminology
Summary
The gist: Multi-sample Direct Preference Optimization (mDPO) and Multi-sample Identity Preference Optimization (mIPO) extend traditional post-training alignment methods by utilizing multi-sample comparisons to capture group-wise characteristics like diversity and bias more effectively than single-sample comparisons, proving superior in optimizing collective model properties across various generative tasks.
Motivation and Problem Statement
Current post-training methods such as reinforcement learning from human feedback (RLHF) and direct alignment from preference methods (DAP) primarily utilize single-sample comparisons, which often fail to capture critical characteristics such as generative diversity and bias, which are more accurately assessed through multiple samples. Evaluating a model’s creativity/consistency or detecting biases requires analyzing the variability and diversity across multiple outputs, not just individual ones. For example, while LLMs are proficient in crafting narratives, they often show limitations in generating a diverse representation of genres (Patel et al., 2024; Wang et al., 2024). Additionally, these models tend to have lower entropy in their predictive distributions after post-training, leading to limited generative diversity (Mohammadi, 2024; Wang et al., 2023; Wiher et al., 2022; Khalifa et al., 2020). Inconsistencies in generation is also a crucial issue that needs to be addressed to make models more reliable (Liu et al., 2013; Bubeck et al., 2013). The aforementioned failures cannot be captured by a single sample; instead, they are distributional issues (see Fig. 1 for an illustration)
Methodology: Multi-sample Optimization
The paper introduces Multi-sample Direct Preference Optimization (mDPO) and Multi-sample Identity Preference Optimization (mIPO), which are extensions of the prior DAP methods–DPO (Rafailov et al., 2024) and IPO (Azar et al., 2024) Unlike their predecessors, which rely on single-sample comparisons, mDPO and mIPO utilize multi-sample comparisons to better capture group-wise or distributional characteristics.
The likelihood of a group Gw preferred over Gl is defined as p(Gw ≽ Gl x) = Φ(r(Gw, x) − r(Gl, x)), where Φ can be a sigmoid function (recovering the Bradley-Terry model) The goal of RLHF is to maximize the reward under reverse KL constraints, which can be captured by an objective like LmDPO = E(x,Gw,Gl)∼D − log σ βEyw∼Gw log πθ(ywx)πref(ywx) − βEyl∼Gl log πθ(ylx)πref(ylx) Similarly, for the IPO variant, the objective is LmIPO = E(x,Gw,Gl)∼DEyw∼Gw log πθ(ywx)πref(ywx) − Eyl∼Gl log πθ(ylx)πref(ylx) − τ−1/22
Experimental Results and Findings
Empirically, the authors demonstrate that multi-sample comparison is more effective in optimizing collective characteristics (e.g., diversity and bias) for generative models than single-sample comparison. In random number generation, both mIPO and mDPO achieve higher uniformity compared to the SFT baseline, with mDPO vs SFT showing a win rate of 0.99 For text-to-image generation, multi-sample optimization enables notable reductions in gender and race biases, with mDPO significantly improving the Simpson Diversity Index over the original SD 1.5 model In creative fiction generation, both quality and diversity are significantly improved; for instance, mIPO (k=5) achieved a win rate of 0.353 against DPO Furthermore, the proposed methods are especially robust against label noise; the gap between win and lose rates narrows as k increases when comparing default labeling versus GPT-4o labeling
Technical Details and Robustness
The paper details a stochastic estimator for efficient optimization, noting that for mIPO, the objective can be expanded to compute an unbiased estimator of the squared difference between expectations, which is equivalent to minimizing l = (Ex∼p[f(x)] − Ex∼q[f(x)] − c)2 The variance of the mini-batch estimator for mIPO is given by Var(ˆl) = O σ 2 p n + σ 2 q m!·σ 2 p n + σ 2 q m + (µp − µq − c) 2!! The analysis of this variance shows that as the sample size k increases, the error of the biased estimator also decreases and performs similarly to the unbiased one The paper concludes that multi-sample comparison is much more advantageous with labeling noise, while single-sample comparison is best with noise-free labels
Conclusion
In this paper, we introduced multi-sample DPO (mDPO) and multi-sample IPO (mIPO), which are novel extensions to the existing DAP methods.
Improvements for AI systems
- Bold Header: Multi-sample Preference Optimization (mDPO/mIPO) for Collective Characteristics
This method allows models to optimize collective characteristics (e.g., diversity and bias) for generative models than single-sample comparison,
improving the alignment of outputs with desired distributional properties by evaluating the model’s performance over distributions of samples rather than individual samples.
- Bold Header: Enhanced Random Number Generation Uniformity
The application of mIPO/mDPO to random number generation demonstrates that these methods can achieve higher uniformity, as shown in Table 1 where mIPO vs IPO
shows a win rate of 0.99 compared to 0.80 for the baseline, leading to a predictive distribution much closer to the uniform distribution.
- Bold Header: Robust Debiasing in Image Generation
For diffusion models, mDPO significantly reduces biases by comparing groups of samples; specifically, Table 2 shows that mDPO improves the Simpson Diversity Index compared to DPO and SFT for both gender and race metrics, resulting in a significantly higher
diversity score.
- Bold Header: Improved Creative Fiction Diversity
In fiction generation, mIPO (k=5) achieves the best results for genre distribution, with a KL-divergence of 0.050 compared to DPO's 0.170, indicating that models finetuned using multi-sample methods exhibit improved diversity at both lexical and semantic levels compared to the baselines.
- Bold Header: Robustness Against Label Noise
The proposed methods are shown to be especially robust against label noise,
as demonstrated by Figure 7, where multi-sample comparison is more robust to labeling noise
than single-sample comparison when using default (noisy) labeling.
Sources
- GPT-4 Technical Report
- Nemotron-4 340B Technical Report
- Qwen Technical Report
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- Training Diffusion Models with Reinforcement Learning
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- MJ-Bench: Is Your Multimodal Reward Model Really a Good Judge for Text-to-Image Generation?
- Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack
- RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
- GPTs are GPTs: An Early Look at the Labor Market Impact Potential of Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- Aligning Language Models with Preferences through f-divergence Minimization
- A Confederacy of Models: a Comprehensive Evaluation of LLMs on Creative Writing
- A Distributional Approach to Controlled Text Generation
- Scalable agent alignment via reward modeling: a research direction
- A Diversity-Promoting Objective Function for Neural Conversation Models
- PopAlign: Population-Level Alignment for Fair Text-to-Image Generation
- Best Practices and Lessons Learned on Synthetic Data
- Trustworthy LLMs: a Survey and Guideline for Evaluating Large Language Models' Alignment
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks