FAMPWQ: Fisher Information-based Adaptive Mixed Precision Weight Quantization for Effective LLM Inference
cs.LG, cs.AI, cs.DC
Submitted: 2026-08-24
Updated: 2026-08-31
Comments: 21 pages, to appear in EMNLP 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on
Terminology
Abstract
Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Qwen Technical Report
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Mistral 7B
- SqueezeLLM: Dense-and-Sparse Quantization
- GPTAQ: Efficient Finetuning-Free Quantization for Asymmetric Calibration
- Fisher Information-based Efficient Curriculum Federated Learning with Large Language Models
- Pointer Sentinel Mixture Models
- Language Models as Knowledge Bases?
- Proximal Policy Optimization Algorithms
- Qwen2 Technical Report
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Model Compression and Efficient Inference for Large Language Models: A Survey
- Beyond Random Missingness: Clinically Rethinking for Healthcare Time Series Imputation
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Qwen3 Technical Report
- MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks