Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Values
Jan Betley, Johannes Treutlein, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans
cs.LG, cs.AI, cs.CR
Submitted: 2026-08-14
Updated: 2026-08-18
Code: https://github.com/TruthfulAI-research/value_
License: http://creativecommons.org/licenses/by/4.0/
The gist: People use language models for practical questions whose answers are difficult to verify.
Terminology
Abstract
People use language models for practical questions whose answers are difficult to verify. We show that models exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user. In one of our evaluations, the user is considering investing in an AI company and wants to know how likely the AI bubble is to pop. Claude Opus 4.8 gives a lower probability when the company under consideration is Anthropic rather than OpenAI. Yet Claude mostly fails to disclose this influence to the user. Covert value leakage is a form of misalignment because it goes against the user's preferences and is likely to mislead them. To investigate this phenomenon, we introduce a suite of evaluations to quantify value leakage and whether models disclose it. We find that models are influenced by different types of values, including preferences for morally good outcomes, for the company that developed them, and for some human leisure activities over others. We often observe large differences among frontier models on the same evaluation. For example, on a Fermi-estimation task, Claude models falsely claim to give unbiased answers in their chain-of-thought, while Qwen models explain how their values bias their answers. Value leakage is a failure mode distinct from sycophancy and reward hacking, and current alignment training and evaluations do not adequately address it.
Sources
- ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Looking Inward: Language Models Can Learn About Themselves by Introspection
- Artificial Influence: An Analysis Of AI-Driven Persuasion
- Reasoning Models Don't Always Say What They Think
- Self-Blinding and Counterfactual Self-Simulation Mitigate Biases and Sycophancy in Large Language Models
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Truthful AI: Developing and governing AI that does not lie
- Monitoring Monitorability
- Counterfactual Simulation Training for Chain-of-Thought Faithfulness
- Robustly Improving LLM Fairness in Realistic Settings via Interpretability
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Frontier Models are Capable of In-context Scheming
- Chunky Post-Training: Data Driven Failures of Generalization
- Frontier Models Can Take Actions at Low Probabilities
- Reasoning Models Will Sometimes Lie About Their Reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks