RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
cs.LG, cs.CL
Submitted: 2026-06-10
Updated: 2026-09-14
Comments: 24 pages, 9 figures, 13 tables
Code: https://github.com/THU-BPM/RLCSD
License: http://creativecommons.org/licenses/by/4.0/
The gist: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified
Terminology
Abstract
On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified solution. However, we show that the resulting distributional gap concentrates on style tokens rather than task-bearing ones, as the hinted model tends to produce shorter, more direct outputs. We term this pathology privilege-induced style drift, which can destabilize training and shorten responses. To address this, we propose RLCSD (Reinforcement Learning with Contrastive on-policy Self-Distillation), which mitigates this drift by contrasting the teacher-student gap under a correct hint against that under a wrong hint, suppressing style shifts induced by hints regardless of correctness and yielding a signal more concentrated on task-bearing tokens. Experiments on Qwen3 (1.7B/4B/8B) and Olmo-3-7B-Think across mathematical and logical reasoning show that RLCSD consistently outperforms GRPO and prior OPSD methods, with additional results on agentic tasks supporting broader applicability. We further show that the contrastive principle is general: it plugs into existing OPSD methods to improve them, and its underlying insight extends to broader cross-model on-policy distillation.
Sources
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Adaptive Teacher Exposure for Self-Distillation in LLM Reasoning
- Self-Policy Distillation via Capability-Selective Subspace Projection
- Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
- DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
- Classifier-Free Diffusion Guidance
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Reinforcement Learning via Self-Distillation
- UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
- Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning
- Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- DeepSeek-V3 Technical Report
- Understanding R1-Zero-Like Training: A Critical Perspective
- Olmo 3
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
- Proximal Policy Optimization Algorithms
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks