Robust Ultra Low-Bit Post-Training Quantization via Stable Diagonal Curvature Estimate
cs.LG
Submitted: 2026-04-15
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging.
Terminology
Abstract
Large Language Models (LLMs) are widely used across many domains, but their scale makes deployment challenging. Post-Training Quantization (PTQ) reduces memory footprint without retraining by leveraging a small calibration set. Recent Hessian-based PTQ methods compensate quantization error via cross-channel dependencies, but such approaches degrade at low bit-widths due to noisy curvature estimates from limited calibration data. We propose DASH-Q, a robust PTQ framework using diagonal Hessian approximation and iterative weighted least squares. By discarding noise-prone dependencies, DASH-Q filters sampling noise while prioritizing the preservation of salient feature power. We outperform other PTQ baselines in ultra low-bit regime, improving zero-shot accuracy by 7.01% on average and up to 14.01% over the strongest baselines across five baseline LLM models, while showing robust and stable performance with very small calibration data.
Sources
- Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
- Calibrating Beyond English: Language Diversity for Better Quantized Multilingual LLM
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Bridging the Gap Between Promise and Performance for Microscaling FP4 Quantization
- High-dimensional sample covariance matrices with Curie-Weiss entries
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- MARLIN: Mixed-Precision Auto-Regressive Parallel Inference on Large Language Models
- Mixtral of Experts
- FlexiQ: Adaptive Mixed-Precision Quantization for Latency/Accuracy Trade-Offs in Deep Neural Networks
- Exploring the Trade-Offs: Quantization Methods, Task Difficulty, and Model Size in Large Language Models From Edge to Giant
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- Pointer Sentinel Mixture Models
- Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering
- SocialIQA: Commonsense Reasoning about Social Interactions
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Qwen3 Technical Report
- RPTQ: Reorder-based Post-training Quantization for Large Language Models
- HellaSwag: Can a Machine Really Finish Your Sentence?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks