STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes
cs.CL
Submitted: 2026-05-13
Updated: 2026-09-20
Comments: Accepted to Findings of EMNLP 2026. Revised camera-ready version. 18 pages, 9 figures, 5 tables. Code available at: https://github.com/chenjux/ECN-STOP
Code: https://github.com/chenjux/ECN-STOP
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long chain-of-thought (Long CoT) reasoning improves performance on multi-step problems, but it also induces overthinking.
Terminology
Abstract
Long chain-of-thought (Long CoT) reasoning improves performance on multi-step problems, but it also induces overthinking. This inefficiency is especially problematic in low-data fine-tuning regimes, where real applications adapt reasoning models with limited supervision and cannot rely on large-scale teacher distillation or heavy test-time control. To address this, we propose STOP (Structured On-policy Pruning), an on-policy algorithm for analyzing and pruning long-form reasoning traces. STOP constructs self-distilled traces from the model. Then it maps each trace into a structured reasoning interface through node segmentation, taxonomy annotation, and reasoning-tree construction. On top of this interface, we introduce ECN (Earliest Correct Node), which retains the shortest prefix ending at the earliest node. Experiments on DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-LLaMA-3-8B across GSM8K, Math 500, and AIME 2024 show that STOP reduces generated tokens by 19.4% to 42.4% while largely preserving accuracy in low-data fine-tuning. Beyond efficiency, our analyses show that STOP induces much smaller distributional shift than teacher-guided pruning, improves the structural efficiency of generated reasoning, and reallocates reasoning effort away from redundant verification and backtracking toward more productive exploration.
Sources
- Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
- Training Verifiers to Solve Math Word Problems
- Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OpenAI o1 System Card
- Style over Substance: Distilled Language Models Reason Via Stylistic Replication
- O1-Pruner: Length-Harmonizing Fine-Tuning for O1-Like Reasoning Pruning
- CODI: Compressing Chain-of-Thought into Continuous Space via Self-Distillation
- Kimi k1.5: Scaling Reinforcement Learning with LLMs
- Thoughts Are All Over the Place: On the Underthinking of o1-Like LLMs
- The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
- TokenSkip: Controllable Chain-of-Thought Compression in LLMs
- Think When You Need: Self-Adaptive Chain-of-Thought Learning
- ShorterBetter: Guiding Reasoning Models to Find Optimal Inference Length for Efficient Reasoning
- NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
- Small Models Struggle to Learn from Strong Reasoners
- Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering