ReCal: Calibrating Structured Pruning for On-Policy Distillation Recovery
cs.CL
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/jianghoucheng/ReCal
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- NIRVANA: Structured Pruning Reimagined for Large Language Model Compression
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- MuCRASP: Multimodal Chain-of-thought Reasoning aware Structured Pruning
- SlimLLM: Accurate Structured Pruning for Large Language Models
- Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization
- Distilling the Knowledge in a Neural Network
- Entropy-Aware On-Policy Distillation of Language Models
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- DistiLLM-2: A Contrastive Approach Boosts the Distillation of LLMs
- Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
- ZipLM: Inference-Aware Structured Pruning of Language Models
- Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe
- ShortGPT: Layers in Large Language Models are More Redundant Than You Expect
- Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- LLM Pruning and Distillation in Practice: The Minitron Approach
- A Simple and Effective Pruning Approach for Large Language Models
- DarwinLM: Evolutionary Structured Pruning of Large Language Models
- The LLM Surgeon
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering