When Do Large Language Models Exhibit Unsolicited Deception?
cs.CL
Submitted: 2025-03-31
Updated: 2026-09-09
Journal ref: Transactions of the Association for Computational Linguistics 2026; 14 1936-1962
DOI: 10.1162/TACL.a.786
Code: https://github.com/0xSMT/llm-2x2-games-deception
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are effective at deceiving when prompted to do so.
Terminology
Abstract
Large Language Models (LLMs) are effective at deceiving when prompted to do so. Models that demonstrate better performance on reasoning tasks are also better at prompted deception. But under what conditions do they deceive without instruction to do so? This study evaluates unsolicited deception produced by LLMs in a preregistered experimental protocol using tools from signaling theory. We evaluated a range of 18 proprietary closed-source and open-source LLMs using modified 2x2 games (in the style of the Prisoner's Dilemma) augmented with a phase in which they can freely communicate to the other agent using unconstrained language. This setup creates an opportunity to misrepresent its actions in conditions that vary in how useful doing so might be towards goal satisfaction. The results indicate that 1) all tested LLMs misrepresent their actions in at least some conditions, 2) they are generally more likely to do so in situations in which deception is beneficial, and 3) models exhibiting better reasoning capacity overall tend to misrepresent at higher rates. Taken together, these results suggest a correlational relationship between model reasoning performance and situational deception, and reveal certain contextual factors that affect whether LLMs will misrepresent actions or not in a novel experimental configuration.
Sources
- Concrete Problems in AI Safety
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- The Llama 3 Herd of Models
- AI Alignment: A Comprehensive Survey
- Mixtral of Experts
- Lies, Damned Lies, and Distributional Language Statistics: Persuasion and Deception with Large Language Models
- People cannot distinguish GPT-4 from a human in a Turing test
- Scaling Laws for Neural Language Models
- Scalable agent alignment via reward modeling: a research direction
- Categorizing Variants of Goodhart's Law
- Hoodwinked: Deception and Cooperation in a Text-Based Game for Language Models
- GPT-4 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering