Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation
cs.CL, cs.LG
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/LengSicong/MMR1
Terminology
Sources
- Reinforcement Learning from Rich Feedback with Distributional DAgger
- Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why
- Qwen3-VL Technical Report
- Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
- VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
- Be My Tutor: On-Policy Co-Distillation for Mutual LLM Improvement via Peer Feedback
- Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation
- $\mathcal{X}$-KD: General Experiential Knowledge Distillation for Large Language Models
- Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
- X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs
- Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
- $\boldsymbol{f}$-OPD: Stabilizing Long-Horizon On-Policy Distillation with Freshness-Aware Control
- KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- Gemma 2: Improving Open Language Models at a Practical Size
- LoRA: Low-Rank Adaptation of Large Language Models
- Solving Quantitative Reasoning Problems with Language Models
- Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
- SafeSteer: Localized On-Policy Distillation for Efficient Safety Alignment
- Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering