Unifying Depth and Width Pruning for LLMs via Binary Knapsack Optimization

arXiv:2608.12953 · cs.CL, cs.LG · Submitted 2026-08-13 · Read on arXiv

Palaash Goel, Ayan Sengupta, Akshay Nambi, Tanmoy Chakraborty

Indian Institute of Technology Delhi · Microsoft Research India

cs.CL, cs.LG

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 29 pages, 5 figures, 17 tables

Code: https://github.com/tatsu-lab/stanford_alpaca

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: SNIPER is a two-stage structured pruning framework for large language models (LLMs) that unifies depth and width pruning via binary knapsack optimization.

Terminology

Summary

SNIPER is a two-stage structured pruning framework for large language models (LLMs) that unifies depth and width pruning via binary knapsack optimization. The paper states: "We present SNIPER, a two-stage structured pruning framework that solves a knapsack optimization over coarse-granularity components to yield conditionally optimal parameter allocations with respect to fixed importance estimates, followed by a fine-grained pruning stage to meet strict budget constraints."

The first stage compresses along the depth-axis by solving a 0/1 knapsack problem over coarse-grained components (transformer layers, attention blocks, and MLP blocks), using dynamic programming to guarantee a conditionally optimal selection. The paper explains: "Unlike existing systems that rely on greedy and myopic pruning decisions, SNIPER utilizes dynamic programming to guarantee a conditionally optimal selection of coarse-grained components while pruning along the depth-axis." The second stage applies width pruning to MLP columns to precisely meet the target compression ratio.

A key contribution is the Compression Ratio Adherence Factor (CRAFT), which quantifies budget fidelity. The paper reports: existing pruners deviate from target compression ratios by up to 33%, SNIPER achieves near-exact adherence with a CRAFT score of 0.98. This addresses the capacity slack problem where existing structured pruners exhibit severe 'capacity slack', with the actual compression ratio deviating from the target by up to 33%.

The paper introduces an iterative importance estimation scheme based on marginal contribution, where the importance v(xi) is quantified as the marginal degradation in prediction fidelity using the formula: v(xi) = ∥Z − Zdrop∥22 − ∥Z − Zretain∥22. This ensures importance scores reflect the effect of jointly pruning multiple components, thereby making the scoring function more reliable and representative.

For the fine-grained stage, SNIPER distributes the pruning budget inversely proportional to layer importance using a softmax over negated importance scores, and selects columns based on sensitivity scores: omegaj = Dj,: · (G:,j ⊙ U:,j) where indices with the lowest scores are pruned.

Evaluations were conducted across four diverse architectures (LLaMA-3.1-8B-Instruct, Qwen3-8B, Phi-4-14B, and GPT-OSS-20B) on 18 tasks spanning five domains (generative performance, world understanding, domain-specific knowledge, NLU & NLI, and safety/bias/ethics). The paper reports: Across all pruning configurations, SNIPER achieves an excellent mean rank of 1.25, indicating its robust cross-architectural generalizability and excellent reliability. SNIPER also demonstrates lower task-level standard deviation, with the paper noting SNIPER reduces task-wise standard deviation by up to 60% relative to prior methods.

On larger architectures, "On Phi-4-14B, SNIPER attains an Avg RP of 81.58%, substantially outperforming ShortGPT and 2SSP by 4.4% and 11.42%, respectively. SNIPER demonstrates similar superiority on GPT-OSS-20B, achieving an Avg RP of 86.99% - 2.31% higher than 2SSP and 16.48% better than ShortGPT."

Ablation studies show both stages are complementary: "Coarse pruning alone reaches a strong Avg RP of 97.47%, but its high Std RP (10.50%) reflects a key limitation... The fine-grained variant attains a marginally better Avg RP (97.95%) with far lower Std RP (2.90%)... The two stages are thus complementary: coarse pruning delivers hardware-friendly speedups at the cost of consistency, while fine-grained pruning restores consistency but lacks efficiency gains due to induced irregularities."

The paper also demonstrates robustness to calibration data, showing SNIPER's pruning strategy is significantly less sensitive to calibration choices. Importance scores are transferable across compression ratios with only marginal degradation in average retention performance. The discretizing factor α is robust across a wide range (8 to 32768) while maintaining CRAFT of 0.98.

The paper concludes: "SNIPER achieves precise budget adherence, strong performance retention, and reduced task-level variance across diverse model architectures... These results underscore the value of non-greedy optimization as a principled foundation for reliable and deployable LLM pruning."

Improvements for AI systems

Improvements to AI Systems:

  1. Optimal Non-Greedy Pruning via Dynamic Programming: Replace greedy, myopic pruning heuristics with a two-stage knapsack-based optimization. The improved AI system can guarantee conditionally optimal selection of coarse-grained components (layers, attention, MLP blocks) under hard budget constraints, eliminating the capacity slack problem where actual compression deviates by up to 33% from targets.

  2. Precise Compression Ratio Adherence (CRAFT): Integrate the Compression Ratio Adherence Factor as a built-in metric and optimization target. The improved system can achieve near-exact adherence (CRAFT = 0.98) to any user-specified compression ratio, making deployment predictable for hardware with strict memory or latency budgets.

  3. Joint Importance Estimation for Reliable Scoring: Use marginal-contribution-based importance scores (v(xi) = ∥Z − Zdrop∥22 − ∥Z − Zretain∥22) that account for joint pruning effects. The improved system can produce more reliable and representative importance estimates, avoiding the pitfalls of isolated scoring that mislead pruning decisions in deep networks.

  4. Calibration-Robust Pruning: Leverage SNIPER's demonstrated insensitivity to calibration data. The improved system can prune effectively without requiring carefully curated calibration sets, reducing engineering overhead and improving robustness in real-world deployment where calibration data is noisy or limited.

  5. Transferable Importance Scores Across Compression Ratios: Use importance scores that remain valid across different target compression levels. The improved system can generate a single importance map and apply it to multiple deployment scenarios (e.g., 20%, 30%, 50% compression) with only marginal performance degradation, enabling rapid A/B testing of model variants.

  6. Reduced Task-Level Variance for Stable Performance: Adopt the two-stage complementary approach (coarse depth + fine-grained width) to lower task-wise standard deviation by up to 60%. The improved system can deliver consistent performance across diverse tasks (generation, knowledge, safety, NLU) rather than excelling on some and failing on others—critical for general-purpose assistants.

  7. Hardware-Friendly Speedup with Consistency Restoration: Use coarse pruning for efficient hardware execution (e.g., layer removal) and fine-grained pruning to restore consistency. The improved system can achieve both high average retention (97.95%) and low variance (2.90% Std RP), balancing efficiency and reliability in production.

  8. Cross-Architectural Generalizability: Apply the framework across diverse model families (LLaMA, Qwen, Phi, GPT-OSS) without architecture-specific tuning. The improved system can prune any transformer-based LLM with a mean rank of 1.25 across 18 tasks, making it a universal pruning tool for model compression pipelines.

  9. Robust Discretization Factor Handling: Incorporate the α-robustness (range 8 to 32768) to avoid sensitivity to hyperparameter choices. The improved system can maintain high budget fidelity without extensive hyperparameter search, simplifying the pruning workflow.

  10. Deployable Pruning for Large-Scale Models: Enable efficient pruning of 14B–20B parameter models with high retention (81–87% Avg RP) and significant speedups over prior methods (up to 16.48% better than ShortGPT). The improved system can compress state-of-the-art LLMs for edge or on-premise deployment without sacrificing task performance.

Sources

Related papers