Length Penalties Make Chain-of-Thought Less Monitorable
cs.AI, cs.CL, cs.LG
Submitted: 2026-07-08
Updated: 2026-09-21
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost.
Terminology
Abstract
Recent work trains reasoning models with length penalties to curb overthinking and cut inference cost. We show that these penalties make the chain of thought less monitorable. A length-compressed model still lets misleading hints steer its answers, but it less often verbalizes their influence. We train Qwen3-4B and Qwen3-14B with reinforcement learning under length penalties targeting 60% down to 30% of baseline chain-of-thought length, then evaluate them with nine types of biasing hints on held-out MMLU-Pro-R and four transfer benchmarks. A chain is faithful when an LLM monitor can tell from it that the hint influenced the answer. At the 30% target, accuracy stays near baseline and wrong-answer hints switch answers as often as before. Yet faithfulness drops on every evaluation set for both models, by 39% for Qwen3-14B and 35% for Qwen3-4B on MMLU-Pro-R. A control trained with the same correctness and format rewards but no length penalty leaves faithfulness intact or raises it. Shortening alone does not explain the drop. Compressed chains mention the hint 7 to 35 percentage points less often than the uncompressed model's chains shortened to the same length by random sentence deletion, across both model sizes and all five evaluation sets. Length penalties therefore trade monitorability for inference cost by removing the evidence monitors depend on.
Sources
- L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
- MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- CoT Red-Handed: Stress Testing Chain-of-Thought Monitoring
- Training Language Models to Reason Efficiently
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs
- Reasoning Models Don't Always Say What They Think
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- Stable Reinforcement Learning for Efficient Reasoning
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Monitoring Monitorability
- ThinkPrune: Pruning Long Chain-of-Thought of LLMs via Reinforcement Learning
- What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
- Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
- DeepCompress: A Dual Reward Strategy for Dynamically Exploring and Compressing Reasoning Chains
- Understanding R1-Zero-Like Training: A Critical Perspective
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection