Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update
cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Accepted to EMNLP 2026 Findings. Code and data: https://github.com/dependentsign/sycophancy-rational-updating
Code: https://github.com/dependentsign/sycophancy-rational-updating
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models often exhibit sycophancy, revising their answers to align with users when users push back.
Terminology
Abstract
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported-Yielding and Rational-Updating. Prior work focuses primarily on suppressing Unsupported-Yielding, while overlooking its effect on Rational-Updating. We address this gap with a two-turn evaluation framework that measures the two behaviors separately. Across representative training-time and inference-time interventions, we find that anti-sycophancy methods often encounter a trade-off in which reducing Unsupported-Yielding can sacrifice Rational-Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone-dependent selectivity gains. Overall, our results suggest that anti-sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational-Updating while reducing Unsupported-Yielding.
Sources
- A Rational Analysis of the Effects of Sycophantic AI
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
- Gemma 3 Technical Report
- The Llama 3 Herd of Models
- How to use and interpret activation patching
- Are You Sure? Challenging LLMs Leads to Performance Drops in The FlipFlop Experiment
- Emotion Concepts and their Function in a Large Language Model
- When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- Simple synthetic data reduces sycophancy in large language models
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering