Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
cs.LG, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/Lossfunk/authority-bias
Terminology
Sources
- Refusal in Language Models Is Mediated by a Single Direction
- PIQA: Reasoning about Physical Commonsense in Natural Language
- PARROT: Persuasion and Agreement Robustness Rating of Output Truth -- A Sycophancy Robustness Benchmark for LLMs
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Training Verifiers to Solve Math Word Problems
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- Ask don't tell: Reducing sycophancy in large language models
- Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
- There Is More to Refusal in Large Language Models than a Single Direction
- A Mechanistic View of Authority Hierarchy in LLM Sycophancy
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Who Endorsed It? Measuring Authority Bias Across Expertise Levels in Language Models
- NRC VAD Lexicon v2: Norms for Valence, Arousal, and Dominance for over 55k English Terms
- Steering Language Models With Activation Engineering
- Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- Simple synthetic data reduces sycophancy in large language models
- Do as We Do, Not as You Think: the Conformity of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks