Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection
cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning,
Terminology
Abstract
Chain-of-thought (CoT) monitoring is a safety strategy where the reasoning of a large language model "actor" is inspected by a "monitor" (often another language model) for signs of unsafe planning, deception, or misalignment. We find that planting harmful but benign-sounding reasoning in the actor's context can steer it to perform adversarial actions while evading monitors, an attack we term "plan injection". We initially discover this attack in the multiple-choice question-answering monitorability setting proposed by Lanham et al. (2023), using the investigator-agent elicitation framework of Li et al. (2025). We generalize the attack and show that the discovered behavior scales to harder tasks (achieving 25-33% monitor evasion rates across different monitorability benchmarks) and larger models such as DeepSeek-R1. Across the settings we study, actor models not only follow injected plans but also paraphrase them as their own reasoning, without explicit attribution to the injections. Finally, we find cases where extra monitor resources cause harm - giving the monitor access to the injected plan drops detection by as much as 50% in the Bio-Math task and in a case study on monitor reasoning budget, we find transcripts where additional thinking tokens are spent rationalizing the injected plan rather than flagging it.
Sources
- Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Reasoning Models Don't Always Say What They Think
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- A Pragmatic Way to Measure Chain-of-Thought Monitorability
- Monitoring Monitorability
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Noticing the Watcher: LLM Agents Can Infer CoT Monitoring from Blocking Feedback
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Reasoning Under Pressure: How do Training Incentives Influence Chain-of-Thought Monitorability?
- OpenAI GPT-5 System Card
- Reasoning Models Struggle to Control their Chains of Thought
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection