Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-15
Updated: 2026-09-15
Comments: 15 pages, 9 tables. Code, prompts, answer keys, and every scored output: https://github.com/IamArmanNikkhah/easy-to-catch-a-liar
Code: https://github.com/IamArmanNikkhah/easy-to-catch-a-liar
License: http://creativecommons.org/licenses/by/4.0/
The gist: An agent that learns from rewards has to trust whatever reports those rewards.
Terminology
Abstract
An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly change, either the world changed or the reporter broke. From the reports alone these are indistinguishable, and reinforcement learning theory shows that no amount of further experience separates them. The prescribed escape is richer data about the reporter itself. We ask whether a frozen language model, handed exactly that data, uses it. We build a two-option game in which a payout swap and a lying reporter produce byte-identical histories. Then we add one verified record: an independent check of one round's real result, printed beside what the reporter said about that round. That single line settles the case. We ask three large models, from two families, to answer one question with one letter. Is the reporter honest or lying? They catch a lying reporter almost perfectly. At the 70B class that holds in every condition we tried; the 32B model slips in one wording. They clear an honest reporter far less often, and how often depends on things that should not matter. Averaged over rounds, letters, and wordings, a 72B model calls an honest reporter a liar 38% of the time when nothing has changed at all, and 58% of the time when the payouts moved. A 70B model from a second family calls an honest reporter a liar 26% and 48% of the time. The failure is not one of reading, because in the situation where nothing changed the same models score 0.96 to 1.00 with the answer printed in the prompt. Which surface feature drives it differs by family. For the Qwen models it is which round the record names, and for Llama it is which letter stands for "honest." Adding the record to a prompt that already states the answer makes Llama less likely to give that answer. We had registered a prediction for that 58% before the run: 35%. The failure is larger than we expected.
Sources
- Reinforcement Learning with a Corrupted Reward Channel
- The Llama 3 Herd of Models
- Stochastic bandits robust to adversarial corruptions
- Inverse Scaling: When Bigger Isn't Better
- Qwen2.5 Technical Report
- LLMs are Greedy Agents: Effects of RL Fine-tuning on Decision-Making Abilities
- Large Language Models Are Not Robust Multiple Choice Selectors
- Cultural Binding Heads in Language Models
- Knowing but Not Correcting: Routine Task Requests Suppress Factual Correction in LLMs
- Bias by Necessity: Impossibility Theorems for Sequential Processing with Convergent AI and Human Validation
- Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks