Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
cs.LG, cs.CL
Submitted: 2026-08-31
Updated: 2026-09-24
Comments: 20 pages, 12 figures
Code: https://github.com/THUDM/slime
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR).
Terminology
Abstract
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base Qwen3-1.7B, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Sources
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- GLM-5: from Vibe Coding to Agentic Engineering
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Where Hindsight Credit Can Reside: A Signed-Capacity View of Token Updates in RLVR
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
- Reinforcement Learning via Self-Distillation
- Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distillation Enables Continual Learning
- Backtracking When It Strays: Mitigating Dual Exposure Biases in LLM Reasoning Distillation
- On the Position Bias of On-Policy Distillation
- Trust Region On-Policy Distillation
- TIP: Token Importance in On-Policy Distillation
- Qwen3 Technical Report
- Self-Distilled RLVR
- Self-Rewarding Language Models
- VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks