On-Policy Self-Distillation without Any Supervision
Yijiang Li, Bingyang Wang, Yijun Liang, Yunjie Tian, Di Fu, Nuno Vasconcelos
cs.LG
Submitted: 2026-08-09
Updated: 2026-08-11
Comments: Project page at https://williamium3000.github.io/u-opsd/ and code at https://github.com/williamium3000/u-opsd
Code: https://github.com/williamium3000/u-opsd
Project page: https://williamium3000.github.io/u-opsd
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs).
Terminology
Abstract
On-policy (Self-)Distillation (OPD / OPSD) has shown strong potential for post-training large language models (LLMs). However, existing methods still rely heavily on external supervision, including ground-truth signals, environmental feedback, or guidance from larger models, and therefore fall short of genuine "self"-distillation. In this study, we show that on-policy self-distillation can be achieved using only a model's own generations via internal consistency. We propose unsupervised on-policy self-distillation (U-OPSD). U-OPSD first samples multiple rollouts and constructs a pseudo solution by majority vote under a self-consistency threshold. It then conditions the model's distribution on the pseudo-solution and distills itself on the disagreeing completions, allowing the model to correct itself precisely where it is confidently wrong. Across diverse benchmarks, base models, and training settings, U-OPSD consistently improves over the base models and matches or surpasses supervised methods with ground truth (GT) such as OPSD and GRPO. On five mathematical reasoning benchmarks, i.e., AIME24, AIME25, HMMT25, MATH500, and AMC23, U-OPSD improves over the base model by 8.5% and 10.7% on Qwen3 non-thinking mode at 4B and 8B scales, and outperforms OPSD by 3.2% and 2.3% on average, respectively. In thinking mode, U-OPSD stays on par with OPSD, ahead by 0.9% at 4B and level at 8B and surpassing GRPO by 0.7% and 1.1%, respectively. Code is available at [https://github.com/williamium3000/u-opsd](https://github.com/williamium3000/u-opsd).
Sources
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- OpenThoughts: Data Recipes for Reasoning Models
- Reinforced Self-Training (ReST) for Language Modeling
- R-Zero: Self-Evolving Reasoning LLM from Zero Data
- Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
- Self-Evolving Visual Questioner
- HERO: Hindsight-Enhanced Reflection from Environment Observations for Agentic Self-Distillation
- When Are Teacher Tokens Reliable? Position-Weighted On-Policy Self-Distillation for Reasoning
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Privileged Information Distillation for Language Models
- Maximizing Confidence Alone Improves Reasoning
- CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
- Can Large Reasoning Models Self-Train?
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Supervised On-Policy Distillation for Reasoning Language Models
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
- Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations
- SDRT: Enhance Vision-Language Models by Self-Distillation with Diverse Reasoning Traces
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- Self-rewarding correction for mathematical reasoning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks