MI-Distillation: Selecting from Model-Interpolated Instruct-Reasoning Data Spectrum for Chain-of-Thought Distillation
cs.CL, cs.AI
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning.
Terminology
Abstract
Recent advances in large reasoning models (LRMs) have shown strong performance on complex problems through long chain-of-thought (Long CoT) reasoning. However, distilling such trajectories into smaller student models remains challenging: direct Long CoT supervision often provides limited gains and can be less effective than concise Short CoT rationales. In this work, we investigate this phenomenon from a gradient-centric perspective. Our analysis shows that Long CoT induces larger gradient magnitudes and more concentrated update directions than Short CoT, with this effect becoming more pronounced as student model capacity increases. These findings suggest that effective Long CoT distillation requires balancing the reasoning information density of reasoning trajectories with their distributional alignment to the student model. Motivated by this insight, we propose Model Interporlation Distillation (MI-Distillation), a framework that constructs a continuous Instruct-Reasoning data spectrum through model interpolation. To select suitable trajectories from this spectrum, we further introduce Sequential Learnable Surprisal Score (SeqLSS), which favors reasoning paths that are both informative and learnable for the student. Extensive experiments on reasoning benchmarks show that MI-Distillation consistently improves small model CoT distillation over strong Long CoT baselines.
Sources
- Stealing Part of a Production Language Model
- Training Verifiers to Solve Math Word Problems
- The Signal is in the Steps: Local Scoring for Reasoning Data Selection
- Efficient Reasoning Models: A Survey
- In Their Own Words: Reasoning Traces Tailored for Small Models Make Them Better Reasoners
- The Llama 3 Herd of Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- Distilling the Knowledge in a Neural Network
- Editing Models with Task Arithmetic
- OpenAI o1 System Card
- What Happened in LLMs Layers when Trained for Fast vs. Slow Thinking: A Gradient Perspective
- When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Kimi K2.5: Visual Agentic Intelligence
- Revisiting Model Interpolation for Efficient Reasoning
- Qwen2.5 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering