The Undetected Damage of Quantization on Retrieval and How to Fix It
cs.LG, cs.AI
Submitted: 2026-09-21
Updated: 2026-09-27
Comments: 5 figures, 4 tables in the main paper
Project page: https://mobiusml.github.io/hqq_blog
License: http://creativecommons.org/licenses/by/4.0/
The gist: We show that a quantized model that keeps its classification accuracy still changes 14 to 46% of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage.
Terminology
Abstract
We show that a quantized model that keeps its classification accuracy still changes 14 to 46% of its top-1 retrieval results, and that aggregate ranking metrics reveal only part of this damage. We tie this failure to the gap between the two highest scores and use that gap to decide when a quantized answer can be trusted and where additional precision should be spent. We show that the top-1 result is guaranteed to survive quantization only when this gap exceeds twice the largest rounding error. In classification, scores are the logits, and the loss function pushes the correct class away from other classes, encouraging this gap. In retrieval, scores are query-document scores, and nothing separates the top-1 item from the second. This gap can be measured without labels. Before deployment, it predicts which models will break under quantization, and at deployment time it tells, per input, whether the quantized answer still matches the full-precision answer. Most classification inputs have a gap wide enough to trust the quantized answer, but few retrieval queries do. That gap motivates a different fix in each task. In retrieval, spending extra bit-width on the layers whose quantization moves the gap most recovers up to three-quarters of an extra bit's benefit for half its cost. In classification, routing the few low-gap inputs to full precision recovers most of the lost accuracy at a fraction of the cost.
Sources
- MS MARCO: A Human Generated MAchine Reading COmprehension Dataset
- Training Dynamics Impact Post-Training Quantization Robustness
- Deep Learning for Classical Japanese Literature
- Detecting a Route Flip Is Easier Than Knowing Whether to Fix It: Causal Route-Mediated Damage in Quantized Mixture-of-Experts
- What Do Compressed Deep Neural Networks Forget?
- Characterising Bias in Compressed Models
- Characterizing and Understanding the Behavior of Quantized Models for Reliable Deployment
- Boundary-Aware Quantization: Finite-Scale Decision Geometry of Neural Classifiers
- Towards General Text Embeddings with Multi-stage Contrastive Learning
- A Practical Mixed Precision Algorithm for Post-Training Quantization
- Value-and-Structure Alignment for Routing-Consistent Quantization of Mixture-of-Experts Models
- Zero-Shot Quantization via Weight-Space Arithmetic
- BEIR: A Heterogenous Benchmark for Zero-shot Evaluation of Information Retrieval Models
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- DoReFa-Net: Training Low Bitwidth Convolutional Neural Networks with Low Bitwidth Gradients
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks