DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation
cs.CL, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/THUDM/slime
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Rethinking On-Policy Distillation of Large Language Models II: One Training Example
- GLM-5: from Vibe Coding to Agentic Engineering
- MiniLLM: On-Policy Distillation of Large Language Models
- OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
- Measuring Mathematical Problem Solving With the MATH Dataset
- The Curious Case of Neural Text Degeneration
- DASH: Divergence-Adaptive Supervision Horizons for On-Policy Self-Distillation of Reasoning Models
- Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level
- When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation
- Bridging Reasoning Trajectories in On-Policy Distillation via Near-Future Guidance
- Entropy-Aware On-Policy Distillation of Language Models
- Kimi K3: Open Frontier Intelligence
- Solving Quantitative Reasoning Problems with Language Models
- CROP: Task Relevance via Counterfactuals for Selective On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering