Playing Devil's Advocate: Off-the-Shelf Persona Vectors Rival Targeted Steering for Sycophancy

arXiv:2605.21006 · cs.AI, cs.CL, cs.LG · Submitted 2026-05-20 · Read on arXiv

cs.AI, cs.CL, cs.LG

Submitted: 2026-05-20

Updated: 2026-09-30

Comments: Spotlight at the 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026. Revised framing and related work; author-order and footnote corrections

Journal ref: 2nd Workshop on Epistemic Intelligence in Machine Learning, ICML 2026 (Spotlight)

License: http://creativecommons.org/licenses/by/4.0/

The gist: Sycophancy is the tendency of language models to agree with users irrespective of correctness.

Terminology

Abstract

Sycophancy is the tendency of language models to agree with users irrespective of correctness. Prior work has extracted sycophancy persona vectors and causally controlled this trait through activation steering (Chen et al., 2025; arXiv:2507.21509). We ask whether existing vectors for general roles, extracted without targeting sycophancy, transfer to this mitigation task. We compare critical and conformist role vectors with a sycophancy-targeted Contrastive Activation Addition (CAA) baseline on a held-out, counterbalanced PhilPapers benchmark, using task-specific coefficient tuning. On Gemma 2 27B and Qwen 3 32B, the selected critical-role vectors achieve mean sycophancy-logit reductions approximately 68% and 98% as large as CAA's, respectively. Conformist-role effects are weak and heterogeneous. Role vectors have low absolute cosine similarity with the measured CAA direction, establishing geometric separation at the intervention layer without identifying distinct downstream mechanisms. These results show that general persona vectors can help mitigate sycophancy in LLMs, even when extracted without sycophancy-specific labels. Code: https://anonymous.4open.science/#!/r/Sycophancy-Steering-9DF0/.

Sources

Related papers