Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach
cs.LG, cs.SY, eess.SY
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/timfduffy/syco-bench
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Incentivizing Truthful Language Models via Peer Elicitation Games
- ELEPHANT: Measuring and understanding social sycophancy in LLMs
- Consensus is Not Verification: Why Crowd Wisdom Strategies Fail for LLM Truthfulness
- Ask don't tell: Reducing sycophancy in large language models
- Peer Truth Serum: Incentives for Crowdsourcing Measurements and Opinions
- SycEval: Evaluating LLM Sycophancy
- Gemma 3 Technical Report
- On The Truthfulness of 'Surprisingly Likely' Responses of Large Language Models
- Localizing Model Behavior with Path Patching
- The Llama 3 Herd of Models
- Measuring Sycophancy of Language Models in Multi-turn Dialogues
- LoRA: Low-Rank Adaptation of Large Language Models
- Language Models (Mostly) Know What They Know
- Scaling Laws for Neural Language Models
- TruthfulQA: Measuring How Models Mimic Human Falsehoods
- Eliciting Informative Text Evaluations with Large Language Models
- Sycophancy in Large Language Models: Causes and Mitigations
- Linear Probe Penalties Reduce LLM Sycophancy
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks