KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-03-02
Updated: 2026-09-07
Comments: 10 pages, 5 figures, 5 tables, code is available at: https://github.com/songmzhang/KDFlow
Code: https://github.com/songmzhang/KDFlow
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as
Terminology
Abstract
Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher logits on each student worker using a frozen copy of the teacher's output head. Furthermore, our framework supports both off-policy and on-policy distillation and incorporates cross-tokenizer algorithms through highly extensible and user-friendly APIs. Experiments show that KDFlow achieves a 1.44 times to 6.36 times speedup over MS-SWIFT in off-policy distillation and a 1.43 times to 1.75 times speedup over verl in on-policy distillation. KDFlow further scales to 64 GPUs, achieving 3.68 times and 2.52 times strong-scaling speedups in two representative model configurations. The code and documentation are publicly available.
Sources
- Towards Cross-Tokenizer Distillation: the Universal Logit Distillation Loss for LLMs
- Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
- MiniLLM: On-Policy Distillation of Large Language Models
- Distilling the Knowledge in a Neural Network
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- Defeating the Training-Inference Mismatch via FP16
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Knowledge Fusion of Large Language Models
- Reinforcement Learning Optimization for Large-Scale Learning: An Efficient and User-Friendly Scaling Library
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
- MiMo-V2-Flash Technical Report
- Qwen3 Technical Report
- A Dual-Space Framework for General Knowledge Distillation of Large Language Models
- Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering