What Does Chain-of-Thought Entropy Measure? A Channel Audit of Scaffolding, Routing, and Content
stat.ML, cs.IT, cs.LG, math.IT
Submitted: 2026-08-26
Updated: 2026-08-26
License: http://creativecommons.org/licenses/by/4.0/
The gist: Entropy over chain-of-thought tokens decides which tokens receive the policy gradient, which get pruned, and whether a run has collapsed, yet each such statistic reads a next-token distribution
Terminology
Abstract
Entropy over chain-of-thought tokens decides which tokens receive the policy gradient, which get pruned, and whether a run has collapsed, yet each such statistic reads a next-token distribution mixing three choices: whether to emit connective scaffolding, which connective, and what the substantive continuation should be. Designating a scaffold vocabulary subset separates the three, exactly, for entropy, Kullback--Leibler divergence, and the first-order entropy velocity of a softmax policy. We prove the raw and content conventions disagree about which position is the larger fork on an explicit open region, and bound answer diversity by the content channel plus a leakage term a measured witness certifies. Across twenty-three configurations the scaffold side carries up to 41% of the raw high-entropy set; on a matched-tokenizer ladder, coupling changes only at the math-corpus step while the scaffold's entropy share keeps growing through distillation; a closed-form forecast from one channel correlation tracks selection retention over a 54-point range to five points, unfitted. On compression, the content convention beats raw surprisal in every cell; an answer-leakage audit then corrects our own headline control: re-fed chains earn a quarter to a half of their accuracy from restated answers, and once stripped, no token scorer beats a random contiguous block.
Sources
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning
- The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Measuring Mathematical Problem Solving With the MATH Dataset
- EPiC: Towards Lossless Speedup for Reasoning Training through Edge-Preserving CoT Condensation
- Where does output diversity collapse in post-training?
- Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation
- On the Invariants of Softmax Attention
- Let's Think Dot by Dot: Hidden Computation in Transformer Language Models
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- On the Limits of Learned Importance Scoring for KV Cache Compression
- The Invisible Leash: Why RLVR May or May Not Escape Its Origin
- SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
- Tailoring Teaching to Aptitude: Direction-Adaptive Self-Distillation for LLM Reasoning
- Shorthand for Thought: Compressing LLM Reasoning via Entropy-Guided Supertokens
Related papers
- Behavior of prediction performance metrics with rare events
- Optimal Estimation of Generic Dynamics by Path-Dependent Neural Jump ODEs
- A Posterior-Dynamics Framework for Imaging Inverse Problems with Pretrained Diffusion Priors
- One Permutation Is All You Need: Fast, Deterministic Feature Importance and Model Stress-Testing
- Online Conformal Prediction for Non-Exchangeable Panel Data
- Deep Time-Series Forecasting in 10 Years: A Survey