One Symptom, Three Levers: A Critical Review of On-Policy Self-Distillation
cs.LG, cs.AI, cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Comments: 30 pages, 4 figures. Survey / critical review
Code: https://github.com/siyan-zhao/OPSD
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- RL for Reasoning by Adaptively Revealing Rationales
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- How to Scale Your EMA
- Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
- Exploring Simple Siamese Representation Learning
- A Brief Overview: On-Policy Self-Distillation In Large Language Models
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Born Again Neural Networks
- Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data
- Bootstrap your own latent: A new approach to self-supervised Learning
- Rethinking Reward Supervision: Rubric-Conditioned Self-Distillation
- MiniLLM: On-Policy Distillation of Large Language Models
- When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
- A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
- Blending Supervised and Reinforcement Fine-Tuning with Prefix Sampling
- Reinforcement Learning via Self-Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks