Component Type, Not Reconstruction Error, Predicts Attention Quantization Sensitivity
cs.LG
Submitted: 2026-08-09
Updated: 2026-08-09
Comments: Accepted as a regular paper at IEEE ICMLA 2026; to appear in the conference proceedings
Code: https://github.com/Kasun-Dewage/What-predicts-Quantization-Sensitivity-2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- QAQ: Quality Adaptive Quantization for LLM KV Cache
- Quantization Error Propagation: Revisiting Layer-Wise Post-Training Quantization
- Towards Superior Quantization Accuracy: A Layer-sensitive Approach
- Precision Where It Matters: A Novel Spike Aware Mixed-Precision Quantization Strategy for LLaMA-based Language Models
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
- The Super Weight in Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks