From Layers to Submodules: Rethinking Granularity in Replacement-Based LLM Compression
cs.CL, cs.AI
Submitted: 2026-06-01
Updated: 2026-08-26
Comments: Accepted at EMNLP Findings 2026
Code: https://github.com/eliacunegatti/SubFit
License: http://creativecommons.org/licenses/by/4.0/
The gist: Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules.
Terminology
Abstract
Post-training compression of Large Language Models (LLMs) removes entire architectural components, either deleting them or replacing them with fitted modules. Existing replacement-based methods share two design constraints: full-layer granularity and contiguous selection. We argue that this is overly restrictive: in fact, redundancy in pretrained transformers is not confined to contiguous regions, nor does it evenly distribute between Attention and FeedForward outputs, implying that different strategies best approximate different submodule types and that removable components need not cluster within contiguous depth ranges. Based on this intuition, we introduce SubFit (Submodule-level Fitted residual replacement), which compresses LLMs at the submodule level: Attention and FeedForward submodules are selected non-contiguously, and each receives its own lightweight fitted residual bypass. SubFit operates post-training and requires only calibration data. Across ten LLMs (five base, five instruction-tuned), five sparsity levels from 12.5% to 37.5%, and four replacement-based baselines, SubFit achieves the best aggregate perplexity-accuracy trade-off across the evaluated sparsity levels, with larger gains under aggressive compression. At 25% sparsity, it retains 84.6% of dense downstream accuracy and incurs 2.42x perplexity degradation, against 81.6% and 4.34x for the strongest baselines, while delivering measurable inference speedup and KV-cache savings. Code is available at https://github.com/eliacunegatti/SubFit.
Sources
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- What Matters in Transformers? Not All Attention is Needed
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- Qwen2.5-Coder Technical Report
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- The Llama 3 Herd of Models
- Pointer Sentinel Mixture Models
- DeltaLLM: Compress LLMs with Low-Rank Deltas between Shared Weights
- GLU Variants Improve Transformer
- Orca: Progressive Learning from Complex Explanation Traces of GPT-4
- SlimPajama-DC: Understanding Data Combinations for LLM Training
- On the Limits of Layer Pruning for Generative Reasoning in Large Language Models
- DarwinLM: Evolutionary Structured Pruning of Large Language Models
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering