Taming CoT Obfuscation in VLMs: From Mechanistic Evidence to Activation Enforcement
cs.AI
Submitted: 2026-09-21
Updated: 2026-09-21
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Reasoning Models Struggle to Control their Chains of Thought
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- Output Supervision Can Obfuscate the Chain of Thought
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Unveiling and Causalizing CoT: A Causal Pespective
- Chain of Thought Prompt Tuning in Vision Language Models
- Monitoring Monitorability
- Counterfactual Simulation Training for Chain-of-Thought Faithfulness
- Training on Documents About Monitoring Leads to CoT Obfuscation
- Reinforcement Learning via Self-Distillation
- Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
- Mechanistic Interpretability for Large Language Model Alignment: Progress, Challenges, and Future Directions
- GPT-4o System Card
- Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
- CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Experiential Reinforcement Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection