Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning
cs.CL, cs.LG
Submitted: 2026-09-03
Updated: 2026-09-03
Comments: Published at COLM 2026
Code: https://github.com/kdu4108/importance-advantagehf.co
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer.
Terminology
Abstract
Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.
Sources
- AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Thought Anchors: Which LLM Reasoning Steps Matter?
- Reasoning Models Don't Always Say What They Think
- Training Verifiers to Solve Math Word Problems
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Monitoring Monitorability
- Towards Long-Horizon Interpretability: Efficient and Faithful Multi-Token Attribution for Reasoning LLMs
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Axiomatic Attribution for Deep Networks
- Dueling Network Architectures for Deep Reinforcement Learning
- StepWiser: Stepwise Generative Judges for Wiser Reasoning
- Qwen3 Technical Report
- Do LLMs Really Need 10+ Thoughts for "Find the Time 1000 Days Later"? Towards Structural Understanding of LLM Overthinking
- ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering