Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization
cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/camilablank/sycophancy-dpo
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy.
Terminology
Abstract
Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycophantic agreement is a well-known failure of model alignment, there is limited understanding of how it emerges from model training. In this work, we demonstrate that sycophantic agreement can emerge as an unintended consequence of widely used contrastive preference optimization objectives. Using the OLMo 3 post-training pipeline, we show that, for various pairs of teacher models across three families, there is a strong correlation between the log-ratio of the teacher model sycophantic agreement rates and the resulting student model sycophantic agreement rate. We further demonstrate that this unintended transfer is not limited to DPO but also occurs across 6 other preference optimization objectives. To understand whether this effect can be attributed to particular training examples, we analyze the preference data and find that the sycophancy signal is diffused across the entire dataset rather than concentrated in a sparse set of examples: each example appears neutral, i.e., there are no explicit instances of sycophantic agreement, and filtering based on probe-based data attribution or logit-linear selection fails to mitigate sycophancy without removing a large portion of the dataset. Overall, our findings suggest that the teacher models used to generate preference data can interact with alignment training objectives in unexpected ways, generalizing to undesirable and potentially harmful behaviors like sycophantic agreement.
Sources
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity
- Weird Generalization and Inductive Backdoors: New Ways to Corrupt LLMs
- Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
- Can You Trust an LLM with Your Life-Changing Decision? An Investigation into AI High-Stakes Responses
- Flattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
- Subliminal Learning: Language models transmit behavioral traits via hidden signals in data
- KTO: Model Alignment as Prospect Theoretic Optimization
- Towards Analyzing and Understanding the Limitations of DPO: A Theoretical Perspective
- The Delta Learning Hypothesis: Preference Tuning on Weak Data can Yield Strong Gains
- The Llama 3 Herd of Models
- Measuring Massive Multitask Language Understanding
- Eliciting Behaviors in Multi-Turn Conversations
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- 2 OLMo 2 Furious
- Olmo 3
- How RLHF Amplifies Sycophancy
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- Simple synthetic data reduces sycophancy in large language models
- Probe-Based Data Attribution: Discovering and Mitigating Undesirable Behaviors in LLM Post-Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks