Unbiased Top- k Estimation for On-Policy Distillation
cs.CL, cs.LG, stat.ML
Submitted: 2026-09-28
Updated: 2026-10-01
Terminology
Sources
- Process Reinforcement through Implicit Rewards
- Revisiting On-Policy Distillation: Empirical Failure Modes and Simple Fixes
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Skywork Open Reasoner 1 Technical Report
- Distilling the Knowledge in a Neural Network
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
- Reinforcement Learning via Self-Distillation
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- KL for a KL: On-Policy Distillation with Control Variate Baseline
- Self-Distillation Enables Continual Learning
- HybridFlow: A Flexible and Efficient RLHF Framework
- A Survey of On-Policy Distillation for Large Language Models
- Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation
- Lightning OPD: Efficient Post-Training for Large Reasoning Models with Offline On-Policy Distillation
- MiMo-V2-Flash Technical Report
- On the Position Bias of On-Policy Distillation
- Escaping the KL Agreement Trap in On-Policy Distillation
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- A Survey on Knowledge Distillation of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering