Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models
cs.CL, cs.AI, cs.CR
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: 33 pages,14 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden.
Terminology
Abstract
The rapid capability gains of frontier language models are widely attributed to improved reasoning abilities, yet this cannot be verified as raw CoT traces in closed-source systems are hidden. By registering a simple custom tool through a standard API feature, we induce frontier models to externalize intermediate reasoning. Because these traces may reflect post-hoc rationalization rather than genuine reasoning, we first evaluate against native CoT on open-source models and extend to closed-source frontier models including GPT-6 Astra. We find that the extracted reasoning matches native reasoning performance and substantially outperforms no-reasoning baselines, across competition mathematics, science, and code generation. We then characterize how frontier models structure their intermediate reasoning. Across token efficiency, reasoning-step types, and induced reasoning trees, we identify systematic differences in how models externalize, compress, and organize reasoning. We find that Astra exhibits token-efficient directed reasoning, selecting a correct trajectory earlier, while resolving elementary steps internally and externalizing only crucial reasoning. These findings provide a behavioral lens on frontier-model reasoning beyond benchmark scores.
Sources
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Hidden Thoughts Are Not Secret: Reasoning Trace Exposure in LLMs
- Reasoning Models Can Be Effective Without Thinking
- Stealing Reasoning Traces from Proprietary LLM APIs
- Humanity's Last Exam
- Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- How to Steal Reasoning Without Reasoning Traces
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering