Control-Token Injection Suppresses Chain-of-Thought and Defeats Reasoning-Based Oversight in Tool-Using Agents
cs.CR
Submitted: 2026-09-23
Updated: 2026-09-23
Code: https://github.com/Usama1002/deleting-the-trace
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- Reasoning Models Struggle to Control their Chains of Thought
- AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Automating Agent Hijacking via Structural Template Injection
- In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Gemma 4 Technical Report
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- PromptArmor: Simple yet Effective Prompt Injection Defenses
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- RTBAS: Defending LLM Agents Against Prompt Injection and Privacy Leakage
- MELON: Provable Defense Against Indirect Prompt Injection Attacks in AI Agents
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs