When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Terminology
Sources
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
- When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- Mismatch Matters: On-Policy Distillation Beyond Token Agreement
- Less is More: Early Stopping Rollout for On-Policy Distillation
- Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation
- Pass the Baton: Trajectory-Relayed On-Policy Distillation
- Trajectory-Refined Distillation
- Kimi K3: Open Frontier Intelligence
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Qwen3 Technical Report
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- JustRL: Scaling a 1.5B LLM with a Simple RL Recipe
- A Survey of On-Policy Distillation for Large Language Models
- Entropy-Aware On-Policy Distillation of Language Models
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- 1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering