A Token is Worth over 1,000 Tokens: Efficient Knowledge Distillation through Low-Rank Clone
cs.CL, cs.AI
Submitted: 2025-05-19
Updated: 2026-09-01
Comments: NeurIPS 2025 Spotlight
Code: https://github.com/CURRENTF/LowRankClone
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models.
Terminology
Abstract
Training high-performing Small Language Models (SLMs) remains costly, even with knowledge distillation and pruning from larger teacher models. Existing work often faces three key challenges: (1) information loss from hard pruning, (2) inefficient alignment of representations, and (3) underutilization of informative activations, particularly from Feed-Forward Networks (FFNs). To address these challenges, we introduce Low-Rank Clone (LRC), an efficient pre-training method that constructs SLMs aspiring to behavioral equivalence with strong teacher models. LRC trains a set of low-rank projection matrices that jointly enable soft pruning by compressing teacher weights, and activation clone by aligning student activations, including FFN signals, with those of the teacher. This unified design maximizes knowledge transfer while removing the need for explicit alignment modules. Extensive experiments with open-source teachers (e.g., Llama-3.2-3B-Instruct, Qwen2.5-3B/7B-Instruct) show that LRC matches or surpasses state-of-the-art models trained on trillions of tokens--while using only 20B tokens, achieving over 1,000x training efficiency. Our codes and model checkpoints are available at https://github.com/CURRENTF/LowRankClone and https://huggingface.co/collections/JitaiHao/low-rank-clone-lrc-6828389e96a93f1d4219dfaf.
Sources
- Phi-4 Technical Report
- GPT-4 Technical Report
- Nemotron-4 340B Technical Report
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- InternLM2 Technical Report
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
- The Llama 3 Herd of Models
- MiniPLM: Knowledge Distillation for Pre-Training Language Models
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- MEFT: Memory-Efficient Fine-Tuning through Sparse Adapter
- Distilling the Knowledge in a Neural Network
- Accelerating Deep Learning by Focusing on the Biggest Losers
- SpinQuant: LLM quantization with learned rotations
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Instruction Tuning with GPT-4
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- GLU Variants Improve Transformer
- TAID: Temporally Adaptive Interpolated Distillation for Efficient Knowledge Transfer in Language Models
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering