ADMM-Q: An Improved Hessian-based Weight Quantizer for Post-Training Quantization of Large Language Models
cs.LG
Submitted: 2026-05-11
Updated: 2026-08-31
License: http://creativecommons.org/licenses/by/4.0/
The gist: Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs).
Terminology
Abstract
Quantization is an effective strategy to reduce the storage and computation footprint of large language models (LLMs). Post-training quantization (PTQ) is a leading approach for compressing LLMs. Popular weight quantization procedures, including GPTQ and RTN, suffer in model utility, especially at aggressive quantization levels (sub-4-bit). We propose ADMM-Q, a novel weight quantization algorithm that considers the layer-wise quantization problem. Our algorithm is based on a combinatorial variant of the Alternating Direction Method of Multipliers (ADMM). Our operator-splitting procedure updates weights continuously to minimize the layer-wise reconstruction error, while gradually enforcing the quantization constraints with convergence guarantees. We propose additional algorithmic enhancements (e.g., penalty scheduling, preconditioning, and a local search post-processing step) to make ADMM-Q efficient at LLM scale. ADMM-Q is modular and can be used as a drop-in replacement for any weight quantizer within existing quantization pipelines: ADMM-Q is fully composable with existing techniques including range clipping, learned or random rotations, and activation scaling. Using ADMM-Q in place of GPTQ on Qwen3-8B, we decrease WikiText-2 perplexity in: (i) the W3A16 weight-only setting (12.85 to 10.06); (ii) the W4A8 SmoothQuant procedure (9.29 to 8.68); and (iii) the W2A4KV4 SpinQuant procedure (66.11 to 19.42).
Sources
- GPT-4 Technical Report
- Careful Selection of Knowledge to solve Open Book Question Answering
- Fast and Effective Weight Update for Pruned Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
- The Llama 3 Herd of Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Gemini: A Family of Highly Capable Multimodal Models
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- A Survey on Recognizing Textual Entailment as an NLP Evaluation
- Code Llama: Open Foundation Models for Code
- SocialIQA: Commonsense Reasoning about Social Interactions
- A Simple and Effective Pruning Approach for Large Language Models
- Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks