Recovering General Capabilities via Uncertainty-Calibrated Multi-Teacher On-Policy Distillation
cs.CL
Submitted: 2026-08-27
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
The gist: Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing.
Terminology
Abstract
Specializing large language models to vertical domains improves domain-specific behavior but often degrades general capabilities such as reasoning, coding, instruction following, and creative writing. We study this domain--general trade-off in Multi-Teacher On-Policy Distillation (MOPD), where a specialized student is supervised on its own sampled trajectories by domain and general teachers. Standard MOPD faces two limitations: ordinary on-policy sampling rarely exposes tokens with large positive teacher--student advantages, while the advantage sign alone does not establish whether the resulting update direction is reliable. We propose uncertainty-calibrated MOPD to address these limitations. Dual-temperature sampling broadens the candidate trajectory pool, and positive-advantage-density filtering selects trajectories with stronger positive learning signals. Centered log-likelihood (CLL) filtering then computes an entropy-calibrated teacher-endorsement score and probabilistically retains token updates according to direction--endorsement consistency. Experiments on role-playing and medical-domain specialization show that our method improves the general-capability average over standard MOPD by 4.73% and 10.84%, respectively, while maintaining vertical-domain performance. Ablations and diagnostic analyses further confirm that the gains do not merely result from a larger rollout budget and that the proposed trajectory- and token-level mechanisms address their intended failure modes.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- MathArena: Evaluating LLMs on Uncontaminated Math Competitions
- Llama-Nemotron: Efficient Reasoning Models
- Scaling Laws for Forgetting during Finetuning with Pretraining Data Injection
- HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
- Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
- A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
- How Abilities in Large Language Models are Affected by Supervised Fine-tuning Data Composition
- Baichuan-M3: Modeling Clinical Inquiry for Reliable Medical Decision-Making
- MiniLLM: On-Policy Distillation of Large Language Models
- Distilling the Knowledge in a Neural Network
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMs
- m1: Unleash the Potential of Test-Time Scaling for Medical Reasoning with Large Language Models
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Entropy-Aware On-Policy Distillation of Language Models
- Explain in Your Own Words: Improving Reasoning via Token-Selective Dual Knowledge Distillation
- Sequence-Level Knowledge Distillation
- Overcoming catastrophic forgetting in neural networks
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- Scaling Reasoning Efficiently via Relaxed On-Policy Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering