Pruning Laws for Large Language Models
cs.CL
Submitted: 2025-04-06
Updated: 2026-08-27
Comments: Accepted at EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly growing memory and compute requirements, which makes
Terminology
Abstract
Scaling up model parameters and training data consistently improves the performance of large language models (LLMs), but at the cost of rapidly growing memory and compute requirements, which makes deployment on resource-limited hardware infeasible. Model pruning, a widely used compression technique, reduces inference costs by removing redundant parameters. However, its impact on downstream performance remains unpredictable and is typically assessed only through costly empirical sweeps. To address this gap, we introduce pruning laws, simple and interpretable scaling relations that connect a pruned LLM's post-pruning performance to its unpruned performance and pruning ratio. Across ten LLMs (1.3B-30B parameters), a 20B mixture-of-experts model, three pruning strategies (unstructured, width, and depth), and eight diverse tasks, we show that pruning laws achieve strong predictive accuracy (average extrapolation error less than 7%), reliably quantify performance degradation, and identify critical pruning thresholds beyond which recovery is infeasible. Moreover, we demonstrate that the functional form transfers across dense and mixture-of-experts architectures, pruning methods, and unseen models in zero-shot and one-shot setups, with task- and method-specific coefficients that vary in interpretable ways. These results provide both researchers and practitioners with a principled framework to select pruning strategies, estimate safe pruning ratios without exhaustive tuning, and deploy LLMs efficiently under real-world compute and latency constraints.
Sources
- DeepSeek-V3 Technical Report
- SliceGPT: Compress Large Language Models by Deleting Rows and Columns
- Scaling Laws Do Not Scale
- Broken Neural Scaling Laws
- The Llama 3 Herd of Models
- P$^2$ Law: Scaling Law for Post-Training After Model Pruning
- LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- Measuring Massive Multitask Language Understanding
- Deep Learning Scaling is Predictable, Empirically
- Training Compute-Optimal Large Language Models
- Scaling Laws for Neural Language Models
- Scaling Laws for Precision
- You Only Prune Once: Designing Calibration-Free Model Compression With Policy Learning
- A Simple and Effective Pruning Approach for Large Language Models
- Will we run out of data? Limits of LLM scaling based on human-generated data
- SVD-LLM: Truncation-aware Singular Value Decomposition for Large Language Model Compression
- Pointer Sentinel Mixture Models
- Scaling Data-Constrained Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering