Deception by Omission: Language Models Knowingly Hide Their Mistakes
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/Lucas-Florin/mistake-honesty-eval
Project page: https://lucas-florin.github.io/mistake-honesty-eval/1
Terminology
Sources
- Evaluating & Reducing Deceptive Dialogue From Language Models with Multi-turn RL
- WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Reasoning Models Don't Always Say What They Think
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Alignment faking in large language models
- Beyond Chains of Thought: Benchmarking Latent-Space Reasoning Abilities in Large Language Models
- Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
- Uncovering Deceptive Tendencies in Language Models: A Simulated Company AI Assistant
- Training LLMs for Honesty via Confessions
- Kimi K2.5: Visual Agentic Intelligence
- Evaluating whether AI models would sabotage AI safety research
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour
- Stealing Reasoning Traces from Proprietary LLM APIs
- From Sycophancy to Deception: A Unified Taxonomy for LLM Spontaneous Misalignment
- Constitutional Black-Box Monitoring for Scheming in LLM Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering