From Concept Alignment to Causal Grounding: An Intervention Test of Chain-of-Thought Faithfulness
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-19
Updated: 2026-09-30
Comments: In submission
Code: https://github.com/TransformerLensOrg/TransformerLens
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning.
Terminology
Abstract
Chain-of-thought (CoT) can sound plausible yet be unfaithful to the model's underlying reasoning. Most prior work probes CoT faithfulness through input--output behavior or input attributions, leaving internal computation largely underexplored. We instead cast faithfulness as internal concept grounding: Does a large language model's (LLM) CoT reasoning engage the same internal concepts that support the LLM's direct prediction, and do the shared concepts causally drive its answer? Encoding a prediction pass and a CoT pass with a single shared sparse autoencoder (SAE), a reliable approximator of the latent concepts LLMs use, makes their internal concepts directly comparable. We introduce three correlational metrics of concept-level alignment and a causal metric, Δp, which ablates the shared concepts and measures the drop in answer probability. Across five LLMs and four datasets, concept alignment is generally high, as indicated by the correlational metrics; yet these only identify which concepts are shared, not how much they causally contribute. Δp fills this gap: causal faithfulness varies substantially with model depth, peaking at mid-to-late layers rather than the final ones, and model scale reshapes the layer-wise profile. Moreover, causally important shared concepts are not always verbalized in the CoT. These dissociations suggest that faithfulness cannot be reliably assessed from surface-level or representational correspondence alone; assessing it requires causal tests of whether the internal concepts underlying a CoT actually drive the model's prediction.
Sources
- Faithfulness vs. Plausibility: On the (Un)Reliability of Explanations from Large Language Models
- Reasoning Models Don't Always Say What They Think
- Are DeepSeek R1 And Other Reasoning Models More Faithful?
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
- The Llama 3 Herd of Models
- Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
- Verbalizable Representations Form a Global Workspace in Language Models
- Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
- Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy
- Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
- Gemma 2: Improving Open Language Models at a Practical Size
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
- Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
- Local Causal Attribution of Chain-of-Thought Reasoning
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering