Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models
cs.CL, cs.AI, cs.HC
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/calisley/receptiveness-sycophancy
License: http://creativecommons.org/licenses/by/4.0/
The gist: A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment.
Terminology
Abstract
A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy are also characteristic of conversational receptiveness, a construct from social psychology shown to improve interactions across disagreement. We argue that this overlap creates a construct-validity problem for social sycophancy evaluations. Using a popular moral-advice dataset, we find that responses classified as more socially sycophantic are also more receptive. Further, increasing the receptiveness of human-written responses---while preserving their substantive conclusions---causes them to be classified as more socially sycophantic. This tight coupling raises the possibility that social sycophancy evaluations inadvertently penalize desirable behavior. In a preregistered experiment comparing substantively equivalent responses, participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors. The same overall pattern persists even among participants who believe the original question asker is in the wrong. Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.
Sources
- BASIL: Bayesian Assessment of Sycophancy in LLMs
- SWAY: A Counterfactual Computational Linguistic Approach to Measuring and Mitigating Sycophancy
- Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- Measuring and mitigating overreliance to build human-compatible AI
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- Reward Bias Substitution: Single-Axis Bias Mitigations Redirect Optimization Pressure
- Sycophantic Praise: Evaluating Excessive Praise in Language Models
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- Simple synthetic data reduces sycophancy in large language models
- What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering