Internalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning
cs.CL, cs.LG
Submitted: 2026-05-30
Updated: 2026-09-12
Comments: Accepted by EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated policies reduce
Terminology
Abstract
Reinforcement learning from verifiable rewards improves the reasoning ability of large language models, but often suffers from entropy collapse, in which increasingly concentrated policies reduce rollout diversity and useful learning signals. Existing remedies either constrain the RL objective (e.g., entropy regularization) or adjust sampling temperature during rollout collection, but these interventions remain external to the model parameters. We propose Temperature-Scaled On-Policy Self-Distillation (TS-OPSD), a lightweight policy reheating method that internalizes the exploratory effect of temperature into model parameters. Starting from an entropy-collapsed RL checkpoint, TS-OPSD constructs a self-teacher by applying high-temperature scaling to the model's own logits, then distills the resulting smoother distribution back into the student. This policy reheating requires no external teacher, privileged data, or additional inference cost. Experiments on Qwen3-4B-Base and Qwen3-8B-Base show that policy reheating yields a stronger initialization for continued RL than both standard continued RL and rollout-level temperature reheating. Further analyses show that TS-OPSD mainly reduces output sharpness while preserving intermediate representations, top candidate sets, and reasoning capability. These results suggest that entropy restoration can serve as a simple post-collapse intervention for extending reasoning-oriented RL.
Sources
- The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
- A Unified Revisit of Temperature in Classification-Based Knowledge Distillation
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- MiniLLM: On-Policy Distillation of Large Language Models
- Distilling the Knowledge in a Neural Network
- Reinforcement Learning via Self-Distillation
- Entropy-Aware On-Policy Distillation of Language Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- On the Role of Temperature Sampling in Test-Time Scaling
- Entropy-Preserving Reinforcement Learning
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- BAPO: Stabilizing Off-Policy Reinforcement Learning for LLMs via Balanced Policy Optimization with Adaptive Clipping
- CRISP: Compressed Reasoning via Iterative Self-Policy Distillation
- Self-Distilled RLVR
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- On-Policy Context Distillation for Language Models
- Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models
- EDT: Improving Large Language Models' Generation by Entropy-based Dynamic Temperature Sampling
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
- SWIFT:A Scalable lightWeight Infrastructure for Fine-Tuning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering