Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression

arXiv:2609.15838 · cs.LG, cs.AI · Submitted 2026-09-14 · Read on arXiv

cs.LG, cs.AI

Submitted: 2026-09-14

Updated: 2026-09-14

Comments: 16 pages, to appear at EMNLP 2026 Findings

License: http://creativecommons.org/licenses/by/4.0/

The gist: Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's

Terminology

Abstract

Per-matrix singular value decomposition (SVD) truncation is Eckart-Young optimal in the whitened Frobenius norm, but errors from independently compressed matrices compound through the block's nonlinear forward pass. Inspired in part by hierarchical variational optimization in quantum many-body methods, we introduce a three-level chain that widens optimization scope from individual matrices to Transformer blocks to the full model: whitened SVD (L1), block-level joint optimization (L2), and end-to-end language-modeling loss refinement (L3), all from 256 calibration sequences, with no instruction or recovery data. On LLaMA-7B at 60% compression, the chain reduces WikiText-2 perplexity from 42.1 to 19.1 to 11.4. The block-level stage acts as a regularizer: skipping it worsens Penn Treebank (PTB) perplexity by 24 points, a gap that additional end-to-end training did not close in our experiments. Perplexity gains hold across 20-80% compression, five architectures up to 13B parameters, and both in-distribution and out-of-distribution benchmarks, though the cross-architecture rows use architecture-specific configurations and the ratio sweep was not run under one common protocol. With more calibration data, skipping the block-level stage becomes competitive, revealing an offline compute--data trade-off. We therefore claim improvements only in perplexity and compression fidelity; downstream accuracy remains well below the dense model.

Sources

Related papers