Zipping the Thought: When and How Compressed Reasoning Data Works in LLM Post-Training
cs.AI, cs.LG
Submitted: 2026-05-27
Updated: 2026-08-31
Comments: Findings of EMNLP
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge.
Terminology
Abstract
Large language models (LLMs) can now solve complex problems through long chain-of-thought (CoT) reasoning, but the trade-off between performance and token cost remains a central challenge. To address this issue, supervised fine-tuning (SFT) often uses compressed reasoning data, where CoT traces are shortened into compact forms. However, the effect of such compressed reasoning data on post-training remains poorly understood. In this paper, we propose a taxonomy of CoT consisting of Explicit CoT, which outputs all operations without aggregation, Composed CoT, which combines multiple operations into a single step, and Implicit CoT, which omits intermediate operations. As a controlled mechanistic study, we construct a synthetic compositional reasoning task that allows controlled variation of difficulty, compression granularity, and data size, and conducted a comprehensive set of experiments across different model families and sizes. Notably, we find that (i) coarser CoT requires more SFT data, (ii) compared with Explicit CoT, Composed CoT and Implicit CoT benefit more from data scaling, while Composed CoT benefits from data repetition and Implicit CoT tends to lead to memorization, (iii) unlike SFT, subsequent reinforcement learning (RL) with verifiable rewards (RLVR) decomposes compressed steps learned during SFT, and (iv) unidirectional CoT ordering shows stronger generalization on longer sequential tasks. Our findings provide implications for CoT design under data resource constraints and offer important insights into the mechanisms of SFT and RL in LLM post-training.
Sources
- A Theory for Emergence of Complex Skills in Language Models
- Metacognitive Reuse: Turning Recurring LLM Reasoning Into Concise Behaviors
- How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
- Shape of Thought: When Distribution Matters More than Correctness in Reasoning Tasks
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Evaluating Large Language Models Trained on Code
- On the Measure of Intelligence
- Training Verifiers to Solve Math Word Problems
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- AlphaMaze: Enhancing Large Language Models' Spatial Intelligence via GRPO
- From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
- Implicit Chain of Thought Reasoning via Knowledge Distillation
- S3-CoT: Self-Sampled Succinct Reasoning Enables Efficient Chain-of-Thought LLMs
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Skill-Targeted Adaptive Training
- Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning
- SEER: Self-Enhancing Chain-of-Thought Compression for Reasoning Models
- OpenAI o1 System Card
- Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-Tuning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection