Hidden Reasoning Must Leak, but Need Not Be Readable: Fundamental Opportunities and Limits for Chain-of-Thought Monitoring
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/ggml-org/llama.cpp
Terminology
Sources
- Analyzing and Improving Chain-of-Thought Monitorability Through Information Theory
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Lessons from Studying Two-Hop Latent Reasoning
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
- Quantifying the Necessity of Chain of Thought through Opaque Serial Depth
- Reasoning Models Don't Always Say What They Think
- Exploring the Cryptographic Limits of Transformer Networks
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- A completely uniform transformer for parity
- SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
- Measuring Faithfulness in Chain-of-Thought Reasoning
- ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
- Transformers Can Do Arithmetic with the Right Embeddings
- Robust Steganography from Large Language Models
- Evaluating Frontier Models for Stealth and Situational Awareness
- Tool Use Enables Undetectable Steganography in Multi-Agent LLM Systems
- Preventing Language Models From Hiding Their Reasoning
- Average-Hard Attention Transformers are Constant-Depth Uniform Threshold Circuits
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks