Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families
cs.IR, cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: 24 pages, 3 figures, 10 tables. Measurements, ledger and analysis code: https://github.com/ThakiCloud/skillret-ptq-measurements
Code: https://github.com/ThakiCloud/skillret-ptq-measurements
License: http://creativecommons.org/licenses/by/4.0/
The gist: Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity,
Terminology
Abstract
Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity, prefer a ranking-aware objective over weight reconstruction -- was carried into LLM quantization largely intact. We test that advice on retrieval embedders directly, quantizing five checkpoints from four architecture families across a grid of bit widths and group sizes, and isolating the embedding, attention and feed-forward blocks at each width. Every heuristic fails to transfer as stated. The embedding table never emerges as the dominant isolated protection priority in any family, despite being the largest tensor in several of them. Module sensitivity does not survive as a transferable ordering: at INT4/g16 the spread between modules is too small to allocate against, at INT3 the ordering becomes family-dependent and joint damage stops being the sum of its parts, and at INT2 comparable reconstruction error accompanies retention ranging from 1.3 to 65.9 percent of full precision. A cheap reconstruction proxy is useful for screening uniform bit widths but substantially less reliable for choosing which tensors to protect; its apparent strength across the whole grid is a range-extension artifact. A distilled 109M student at INT3 holds 78.04 NDCG@10 in 68.4 MB and dominates the extreme-PTQ arm of its own 0.6B teacher, 297.9 MB at 64.46, on both size and quality -- but only inside the task it was distilled for. Sizes are byte counts of files that exist rather than arithmetic estimates, and the measurement repository carries the byte provenance for every one of them.
Sources
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs
- BinaryBERT: Pushing the Limit of BERT Quantization
- M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural Networks
- HAWQ: Hessian AWare Quantization of Neural Networks with Mixed-Precision
- Extreme Compression of Large Language Models via Additive Quantization
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference
- SkillRet: A Large-Scale Benchmark for Skill Retrieval in LLM Agents
- I-BERT: Integer-only BERT Quantization
- SqueezeLLM: Dense-and-Sparse Quantization
- When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
- BitNet Text Embeddings
- AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
- SpinQuant: LLM quantization with learned rotations
- The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
- Up or Down? Adaptive Rounding for Post-Training Quantization
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG