Learn from Whoever Is Right: Answer-Verified Multi-Teacher Distillation for Multi-Domain LLMs
cs.LG, cs.AI
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/hexixiang/MT-SDPO
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Advancing General-Purpose Reasoning Models with Modular Gradient Surgery
- DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models
- SciKnowEval: Evaluating Multi-level Scientific Knowledge of Large Language Models
- The Llama 3 Herd of Models
- Distilling the Knowledge in a Neural Network
- Reinforcement Learning via Self-Distillation
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
- WebGPT: Browser-assisted question-answering with human feedback
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Expanding the Capabilities of Reinforcement Learning via Text Feedback
- MiMo-V2-Flash Technical Report
- Olmo 3
- Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
- Qwen3 Technical Report
- Disentangling Task Conflicts in Multi-Task LoRA via Orthogonal Gradient Projection
- Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks