RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

arXiv:2606.11709 · cs.LG, cs.CL · Submitted 2026-06-10 · Read on arXiv

cs.LG, cs.CL

Submitted: 2026-06-10

Updated: 2026-09-14

Comments: 24 pages, 9 figures, 13 tables

Code: https://github.com/THU-BPM/RLCSD

License: http://creativecommons.org/licenses/by/4.0/

The gist: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified

Terminology

Abstract

On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with that under privileged context, typically a verified solution. However, we show that the resulting distributional gap concentrates on style tokens rather than task-bearing ones, as the hinted model tends to produce shorter, more direct outputs. We term this pathology privilege-induced style drift, which can destabilize training and shorten responses. To address this, we propose RLCSD (Reinforcement Learning with Contrastive on-policy Self-Distillation), which mitigates this drift by contrasting the teacher-student gap under a correct hint against that under a wrong hint, suppressing style shifts induced by hints regardless of correctness and yielding a signal more concentrated on task-bearing tokens. Experiments on Qwen3 (1.7B/4B/8B) and Olmo-3-7B-Think across mathematical and logical reasoning show that RLCSD consistently outperforms GRPO and prior OPSD methods, with additional results on agentic tasks supporting broader applicability. We further show that the contrastive principle is general: it plugs into existing OPSD methods to improve them, and its underlying insight extends to broader cross-model on-policy distillation.

Sources

Related papers