Quantization Error Is Spectrally Flat: A Single Random Probe Is a Calibrated, Data-Free Sensitivity Estimator, with Application to Budget-Targeted Mixed-Precision Quantization
cs.CL, cs.AI, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-29
Code: https://github.com/mobiusml/hqq
Terminology
Sources
- KurTail : Kurtosis-based LLM Quantization
- Cherry on Top: Parameter Heterogeneity and Quantization in Large Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Extreme Compression of Large Language Models via Additive Quantization
- Mixture Compressor for Mixture-of-Experts LLMs Gains More
- SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models
- Fidelity Is Not Safety: Gently-Compressed LLMs Pass Every Data-Free Quality Guard Yet Invent Procedure Steps in Agentic Execution
- Q-Strata: Hierarchical Bit Allocation for Mixed-Precision Quantization of Mixture-of-Experts LLMs
- Data-free mixed-precision quantization using novel sensitivity metric
- Pushing the Limits of Large Language Model Quantization via the Linearity Theorem
- MixQuant: Adaptive Mixed-Precision Quantization for Large Language Models
- AlphaQ: Calibration-Free Bit Allocation for Mixture-of-Experts Quantization
- GAMMA: Global Bit Allocation for Mixed-Precision Models under Arbitrary Budgets
- Beyond Scalar Sensitivity: Activation-Aware Mixed-Precision LLM Quantization with Cross-Layer Refinement
- LQER: Low-Rank Quantization Error Reconstruction for LLMs
- Beyond Outliers: A Data-Free Layer-wise Mixed-Precision Quantization Approach Driven by Numerical and Structural Dual-Sensitivity
- BitsMoE: Cost-Aware Bit Allocation in Spectral Space for MoE LLM Quantization
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering