Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression
cs.LG, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 16 pages, to appear at EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's
Terminology
Abstract
Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened SVD (L1), block-level joint optimization (L2), and end-to-end language-modeling loss refinement (L3), all from 256 calibration sequences, with no instruction or recovery data. On LLaMA-7B at 60% compression, the chain reduces WikiText-2 perplexity from 42.1 to 19.1 to 11.4. The block-level stage acts as a regularizer: skipping it worsens Penn Treebank (PTB) perplexity by 24 points, a gap that additional end-to-end training did not close in our experiments. Perplexity gains hold across 20-80% compression, five architectures up to 13B parameters, and both in-distribution and out-of-distribution benchmarks, though the cross-architecture rows use architecture-specific configurations and the ratio sweep was not run under one common protocol. With more calibration data, skipping the block-level stage becomes competitive, revealing an offline compute--data trade-off. We therefore claim improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.
Sources
- Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
- ERC-SVD: Error-Controlled SVD for Large Language Model Compression
- DipSVD: Dual-importance Protected SVD for Efficient LLM Compression
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Distilling the Knowledge in a Neural Network
- LoRA: Low-Rank Adaptation of Large Language Models
- SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
- AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
- Speeding up Convolutional Neural Networks with Low Rank Expansions
- Lillama: Large Language Models Compression via Low-Rank Feature Distillation
- Mistral 7B
- LLaMA: Open and Efficient Foundation Language Models
- IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
- PHF: Privileged Hidden Flow for On-Policy Self-Distillation
- Pointer Sentinel Mixture Models
- Swift-SVD: Theoretical Optimality Meets Practical Efficiency in Low-Rank LLM Compression
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- ARA: Adaptive Rank Allocation for Efficient Large Language Model SVD Compression
- ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
- SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks