The Company You Keep: How LLMs Respond to Dark Triad Traits
cs.CL
Submitted: 2026-03-04
Updated: 2026-09-07
License: http://creativecommons.org/licenses/by/4.0/
The gist: LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy.
Terminology
Abstract
LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy. This pattern may become problematic when interacting with user prompts that reflect negative social tendencies, risking the amplification of harmful behavior. We examine how LLMs respond to user prompts expressing varying degrees of Dark Triad traits (Machiavellianism, Narcissism, and Psychopathy) using a curated dataset. Our analysis reveals systematic differences across models: while all models predominantly exhibit corrective behavior, some generate reinforcing or ambivalent output. Model behavior further varies with severity level and response sentiment. These findings highlight the need for safer conversational systems that can reliably detect and respond to users escalating from benign to harmful requests.
Sources
- The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs
- SycEval: Evaluating LLM Sycophancy
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- The Llama 3 Herd of Models
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Alignment Without Understanding: A Message- and Conversation-Centered Approach to Understanding AI Sycophancy
- Towards Understanding Sycophancy in Language Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- The Dark Patterns of Personalized Persuasion in Large Language Models: Exposing Persuasive Linguistic Features for Big Five Personality Traits in LLMs Responses
- Introducing v0.5 of the AI Safety Benchmark from MLCommons
- When Truth Is Overridden: Uncovering the Internal Origins of Sycophancy in Large Language Models
- Simple synthetic data reduces sycophancy in large language models
- When Large Language Models contradict humans? Large Language Models' Sycophantic Behaviour
- Qwen3 Technical Report
- Exploring the Personality Traits of LLMs through Latent Features Steering
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering