OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead.
Terminology
Abstract
Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, pruning protects weights by statistical salience rather than by their contribution to correct reasoning, so weights behind erroneous computation survive as readily as those behind correct computation. These erroneous patterns then get carried into the pruned model, degrading reasoning quality, producing both lower accuracy and longer reasoning traces. We propose Outcome-Based Calibration for Large Reasoning Model Pruning (OBC-Prune) to close this gap. OBC first constructs difficulty-matched pairs of correct and incorrect rollouts from problems the model answers inconsistently. It then estimates the causal importance of each reasoning sentence through intervention-based analysis, quantifying how removing its influence affects subsequent predictions. These causal importance scores are converted into per-token weights that rescale the calibration activations used by one-shot pruning methods (SparseGPT, Wanda, ALPS), without modifying the underlying pruning algorithms. Experiments on DeepSeek-R1-Distill-Qwen 1.5B, 7B, and 14B models at 40% and 50% sparsity demonstrate consistent improvements over state-of-the-art calibration baselines across most model sizes and sparsity levels on MATH500, LiveCodeBench, and AIME 2025. These results indicate that preserving causally important reasoning circuits is a substantially more effective pruning objective than uniformly preserving observed activations.
Sources
- Thought Anchors: Which LLM Reasoning Steps Matter?
- From LLMs to LRMs: Rethinking Pruning for Reasoning-Centric Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and Pruning
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Beware of Calibration Data for Pruning Large Language Models
- Quantized Reasoning Models Think They Need to Think Longer, but They Do Not
- Reasoning Models Can be Accurately Pruned Via Chain-of-Thought Reconstruction
- LLM-Pruner: On the Structural Pruning of Large Language Models
- Thinking Sparks!: Emergent Attention Heads in Reasoning Models During Post Training
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Movement Pruning: Adaptive Sparsity by Fine-Tuning
- Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
- Think Before You Prune: Self-Reflective Structured Pruning for Reasoning Language Models
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning
- Think Before You Prune: Selective Self-Generated Calibration for Pruning Large Reasoning Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection