Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy
cs.AI, cs.CL, cs.LG
Submitted: 2026-05-20
Updated: 2026-09-30
Comments: Spotlight at the 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026. Revised framing and related work; author-order and footnote corrections
Journal ref: 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026 (Spotlight)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Sycophancy is the tendency of language models to agree with users irrespective of correctness.
Terminology
Abstract
Sycophancy is the tendency of language models to agree with users irrespective of correctness. Prior work has extracted sycophancy persona vectors and causally controlled this trait through activation steering (Chen et al., 2025; arXiv:2507.21509). We ask whether existing vectors for general roles, extracted without targeting sycophancy, transfer to this mitigation task. We compare critical and conformist role vectors with a sycophancy-targeted Contrastive Activation Addition (CAA) baseline on a held-out, counterbalanced PhilPapers benchmark, using task-specific coefficient tuning. On Gemma 2 27B and Qwen 3 32B, the selected critical-role vectors achieve mean sycophancy-logit reductions approximately 68% and 98% as large as CAA's, respectively. Conformist-role effects are weak and heterogeneous. Role vectors have low absolute cosine similarity with the measured CAA direction, establishing geometric separation at the intervention layer without identifying distinct downstream mechanisms. These results show that general persona vectors can help mitigate sycophancy in LLMs, even when extracted without sycophancy-specific labels. Code: https://anonymous.4open.science/#!/r/Sycophancy-Steering-9DF0/.
Sources
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
- BILLY: Steering Large Language Models via Merging Persona Vectors for Creative Generation
- SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
- Discovering Language Model Behaviors with Model-Written Evaluations
- Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors
- Qwen3 Technical Report
- PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
- Steering Llama 2 via Contrastive Activation Addition
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models
- Towards Understanding Sycophancy in Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- Steering Language Models With Activation Engineering
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- Depth-Wise Activation Steering for Honest Language Models
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- Persona Features Control Emergent Misalignment
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection