LowRankArena: A Standardized Evaluation Platform for SVD-Based LLM Compression
cs.CL, cs.LG
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/Zishan-Shao/lowrankarena
License: http://creativecommons.org/licenses/by/4.0/
The gist: SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs).
Terminology
Abstract
SVD-based low-rank compression has become a fast-growing direction for reducing the memory and computational cost of large language models (LLMs). However, meaningful comparison across existing studies remains difficult as prior evaluations use varied benchmarks, inconsistent ratios, and diverse setups, often failing to isolate low-rank effects from auxiliary techniques. As a result, it remains unclear whether reported gains reflect method-level improvements or differences in evaluation protocol. This lack of comparability highlights the need for a unified, reproducible evaluation platform. To address this problem, we present LowRankArena, a standardized evaluation platform for SVD-based LLM compression. LowRankArena unifies task versions, uniform-precision compression budgets, comparison regimes, and inference measurements, and provides a reproducible pipeline with over 3 TiB released compressed checkpoints. Using LowRankArena, our aligned audit of five representative SVD methods reveals that prior findings are highly conditional under standardized protocols: clear leaders and performance tiers shift across backbones and keep ratios, multiple-choice accuracy can hide large perplexity degradation, and nominal low-rank savings yield workload-dependent and often limited end-to-end speedups. Our code is available at: https://github.com/Zishan-Shao/lowrankarena.git.
Sources
- Zero Sum SVD: Balancing Loss Sensitivity for Low Rank LLM Compression
- Lessons from the Trenches on Reproducible Evaluation of Language Models
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- The Llama 3 Herd of Models
- Language model compression with weighted low-rank factorization
- SAES-SVD: Self-Adaptive Suppression of Accumulated and Local Errors for SVD-based LLM Compression
- Pointer Sentinel Mixture Models
- Layer-wise dynamic rank for compressing large language models
- AA-SVD : Anchored and Adaptive SVD for Large Language Model Compression
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- Qwen3 Technical Report
- ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models
- Instruction-Following Evaluation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering