D cubed-MOPD: Dynamic Domain ScheDuling for Efficient Multi-Teacher Distillation
cs.LG, cs.AI
Submitted: 2026-08-25
Updated: 2026-09-16
Code: https://github.com/THUDM/slime
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts.
Terminology
Abstract
Multi-teacher on-policy distillation (MOPD) distills several domain-expert teachers into a single student by minimizing per-domain reverse-KL divergence on the student's own rollouts. Existing approaches typically fix the per-domain data mixture before training, overlooking the fact that different domains converge at substantially different rates: some plateau early while others continue to improve throughout the training budget. A fixed mixture therefore wastes compute on fast-converging domains and undertrains slower-converging ones. To address this, we propose D cubed-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online. Running asynchronously outside the training process, an off-process watcher periodically tracks each domain's KL trajectory, estimates remaining headroom and current improvement rate, and accordingly adjusts the domain sampling ratios without altering the core training loop. Our D cubed-MOPD scales naturally to arbitrary numbers of domains, and the expected benefit grows as more domains introduce more diverse convergence patterns for the scheduler to exploit. On a Qwen3.6-35B-A3B student distilled from four domain-expert teachers, D cubed-MOPD closes 97% of the average student-to-teacher performance gap, compared with 63% for vanilla MOPD, reaches the same peak performance with an approximately 3 times reduction in rollout steps, and surpasses the specialist teachers on three of seven benchmarks.
Sources
- MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
- Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
- Flow-OPD: On-Policy Distillation for Flow Matching Models
- Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe
- Entropy-Aware On-Policy Distillation of Language Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- Filter, Then Reweight: Rethinking Optimization Granularity in On-Policy Distillation
- MiMo-V2-Flash Technical Report
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- When Top-K Misses the Decision: Tool-Call Drift in Multi-Teacher On-Policy Distillation
- A Survey of On-Policy Distillation for Large Language Models
- Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making
- OJBench: A Competition Level Code Benchmark For Large Language Models
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
- H-OPD: Confidence Aware Heterogeneous Multi-Teacher Multimodal On-policy Distillation
- GLM-5: from Vibe Coding to Agentic Engineering
- Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks