Benford's Law as a Distributional Prior for Post-Training Quantization of Large Language Models
cs.LG
Submitted: 2026-01-29
Updated: 2026-09-10
Comments: Accepted paper in Transactions on Machine Learning Research (TMLR). Link to OpenReview: https://openreview.net/forum?id=YiLcQY4Nje
Code: https://github.com/ufopcsilab/benford-quant
License: http://creativecommons.org/licenses/by/4.0/
The gist: Post-training quantization (PTQ) is a practical way to reduce the memory footprint of large language models, but low-bit quantization is sensitive to mismatches between the quantization codebook and
Terminology
Abstract
Post-training quantization (PTQ) is a practical way to reduce the memory footprint of large language models, but low-bit quantization is sensitive to mismatches between the quantization codebook and the empirical weight/activation distributions. We revisit Benford-like leading-digit statistics as a lightweight diagnostic of scale-broad behavior in transformer tensors. Across several model families, we observe a consistent functional dichotomy: transformational nn.Linear weights tend to be Benford-like, whereas LayerNorm parameters systematically deviate. Motivated by this observation, we propose BenQ, a data-free PTQ codebook that uses a simple log-spaced grid as a proxy for scale-broad distributions and applies it selectively to transformational layers while keeping stability-critical parameters in higher precision. In 4-bit group-wise PTQ, BenQ consistently improves over uniform RTN and trades wins with NF4 across architectures and tasks, while remaining substantially simpler than optimization-based methods. We additionally report dynamic activation quantization as an exploratory stress test: the results show that log-spaced grids can reduce RTN failures in some families, but also reveal that outlier handling remains essential for reliable low-bit activation PTQ. Code is available at https://github.com/ufopcsilab/benford-quant.
Sources
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Mistral 7B
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding
- Convolutional Neural Networks using Logarithmic Data Representation
- GANQ: GPU-Adaptive Non-Uniform Quantization for Large Language Models
- Rethinking Neural Networks With Benford's Law
- Gemma: Open Models Based on Gemini Research and Technology
- Qwen Technical Report
- The Llama 3 Herd of Models
- OPT: Open Pre-trained Transformer Language Models
- BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
- Pointer Sentinel Mixture Models
- AFPQ: Asymmetric Floating Point Quantization for LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks