When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
cs.CL, cs.AI
Submitted: 2026-03-27
Updated: 2026-09-17
Comments: 13 pages, 4 figures, 4 tables
Code: https://github.com/modelscope/evalscope
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs.
Terminology
Abstract
Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality generation in distilled models requires careful joint design of both the student architecture and the distillation process. Many prior distillation works evaluate downstream multiple-choice benchmarks by ranking candidate answers with log-likelihood rather than requiring autoregressive generation, which can obscure important differences in model quality. For example, on overlapping benchmarks, we show that a 7B distilled model that nearly matches its teacher to within 0.2 pp under log-likelihood scoring falls behind by 20.8 pp when it must generate answers autoregressively. We investigate this phenomenon with GenDistill, a multi-stage pipeline we designed for distilling a pretrained Transformer into an efficient Hybrid Kimi Delta Attention (Hybrid-KDA) student. Using it as a controlled testbed on Qwen3-0.6B, we systematically ablate six design axes (training objective, loss masking, training duration, dataset selection, parameter freezing, and architecture choice) and evaluate every choice under both log-likelihood and generation-based protocols. We find that log-likelihood-based evaluation consistently underestimates the gap between teacher and student, and can in some cases reverse the ranking of design choices, so conclusions drawn from perplexity-only evaluation may be misleading. Among the factors we study, dataset selection, completion-only masking, and freezing attention layers during post-training have the largest impact on generation quality. Our best distillation recipe, using a Hybrid-KDA model as the student, retains 86-90% of teacher accuracy on knowledge benchmarks while reducing KV cache memory by up to 75% and improving time-to-first-token by 2-4x at 128K-token contexts.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Mitigating Copy Bias in In-Context Learning through Neuron Pruning
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Llamba: Scaling Distilled Recurrent Models for Efficient Language Processing
- Understanding the Skill Gap in Recurrent Language Models: The Role of the Gather-and-Aggregate Mechanism
- DiJiang: Efficient Large Language Models through Compact Kernelization
- Evaluating Large Language Models Trained on Code
- Hybrid Linear Attention Done Right: Efficient Distillation and Effective Architectures for Extremely Long Contexts
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- Are We Done with MMLU?
- RADLADS: Rapid Attention Distillation to Linear Attention Decoders at Scale
- Mamba: Linear-Time Sequence Modeling with Selective State Spaces
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- FlashEVA: Accelerating LLM inference via Efficient Attention
- CMMLU: Measuring massive multitask language understanding in Chinese
- Distilling to Hybrid Attention Models via KL-Guided Layer Selection
- Jamba: A Hybrid Transformer-Mamba Language Model
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering