STOP: Structured On-Policy Pruning of Long-Form Reasoning in Low-Data Regimes

arXiv:2605.13165 · cs.CL · Submitted 2026-05-13 · Read on arXiv

cs.CL

Submitted: 2026-05-13

Updated: 2026-09-20

Comments: Accepted to Findings of EMNLP 2026. Revised camera-ready version. 18 pages, 9 figures, 5 tables. Code available at: https://github.com/chenjux/ECN-STOP

Code: https://github.com/chenjux/ECN-STOP

License: http://creativecommons.org/licenses/by/4.0/

The gist: Long chain-of-thought (Long CoT) reasoning improves performance on multi-step problems, but it also induces overthinking.

Terminology

Abstract

Long chain-of-thought (Long CoT) reasoning improves performance on multi-step problems, but it also induces overthinking. This inefficiency is especially problematic in low-data fine-tuning regimes, where real applications adapt reasoning models with limited supervision and cannot rely on large-scale teacher distillation or heavy test-time control. To address this, we propose STOP (Structured On-policy Pruning), an on-policy algorithm for analyzing and pruning long-form reasoning traces. STOP constructs self-distilled traces from the model. Then it maps each trace into a structured reasoning interface through node segmentation, taxonomy annotation, and reasoning-tree construction. On top of this interface, we introduce ECN (Earliest Correct Node), which retains the shortest prefix ending at the earliest node. Experiments on DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-LLaMA-3-8B across GSM8K, Math 500, and AIME 2024 show that STOP reduces generated tokens by 19.4% to 42.4% while largely preserving accuracy in low-data fine-tuning. Beyond efficiency, our analyses show that STOP induces much smaller distributional shift than teacher-guided pruning, improves the structural efficiency of generated reasoning, and reallocates reasoning effort away from redundant verification and backtracking toward more productive exploration.

Sources

Related papers