Multi-Bitwidth Quantization for LLMs Using Additive Codebooks
cs.LG, cs.CL, cs.IT, math.IT
Submitted: 2026-06-11
Updated: 2026-09-26
Code: https://github.com/chutianxiang/QuIPforall
Terminology
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Extreme Compression of Large Language Models via Additive Quantization
- Remote Inference over Dynamic Links via Adaptive Rate Deep Task-Oriented Vector Quantization
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- The Llama 3 Herd of Models
- Residual Quantization with Implicit Neural Codebooks
- Mistral 7B
- Pointer Sentinel Mixture Models
- Matryoshka Quantization
- AnyBCQ: Hardware Efficient Flexible Binary-Coded Quantization for Multi-Precision LLMs
- Any-Precision LLM: Low-Cost Deployment of Multiple, Different-Sized LLMs
- ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals
- Gemma: Open Models Based on Gemini Research and Technology
- Qwen2.5 Technical Report
- Rethinking LLM Advancement: Compute-Dependent and Independent Paths to Progress
- LLaMA: Open and Efficient Foundation Language Models
- Qinco2: Vector Compression and Search with Improved Implicit Neural Codebooks
- Are We There Yet? A Measurement Study of Efficiency for LLM Applications on Mobile Devices
- A Review on Edge Large Language Models: Design, Execution, and Applications
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks