Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
cs.AI, cs.CL
Submitted: 2026-05-07
Updated: 2026-09-26
Terminology
Sources
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- Reinforcement Learning via Self-Distillation
- On-Policy Context Distillation for Language Models
- CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DAPO: An Open-Source LLM Reinforcement Learning System at Scale
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Solving math word problems with process- and outcome-based feedback
- Let's Verify Step by Step
- AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Distilling the Knowledge in a Neural Network
- A General Language Assistant as a Laboratory for Alignment
- Self-Distillation Enables Continual Learning
- Privileged Information Distillation for Language Models
- GATES: Self-Distillation under Privileged Context with Consensus Gating
- OpenClaw-RL: Train Any Agent Simply by Talking
- HDPO: Hybrid Distillation Policy Optimization via Privileged Self-Distillation
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection