PRQuant: Permutation Residual Quantization for Low-Overhead Inference
cs.LG, cs.CV
Submitted: 2026-08-16
Updated: 2026-09-26
Code: https://github.com/open-compass/opencompass
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers.
Terminology
Abstract
Accuracy of Low-bit quantization of linear layers is often dominated by a small number of outliers. Although existing methods, such as smoothing, rotation, or residual-based approaches, may mitigate this problem, they often introduce new accuracy bottlenecks to weights. Besides, most of these techniques are implemented as online approaches, which can result in heavy execution overheads. To address the afore-mentioned issues, We propose PRQuant (Permutation Residual Quantization), a training-free and low-overhead framework that combines channel reorganization with static weight-side residual compensation. After AWQ-style scaling, PRQuant identifies the input channels that contribute most to weight quantization error, permutes them into contiguous tail blocks, and constructs their residual weight sub-tensors offline. During inference, this contiguous structure enables the activation side to use tail blocks seamlessly without the expensive online gathering operation, and turns scattered residual compensation into a regular tail-augmented GEMM, substantially reducing latency. Experiments demonstrate that PRQuant effectively reduces down-projection reconstruction error. Ablation studies confirm that smoothing and residual compensation are the primary drivers of numerical improvement, while permutation provides a consistent marginal numerical benefit and, more importantly, enables a hardware-friendly contiguous layout that eliminates dynamic gathering overhead. Overall, PRQuant outperforms default MXFP4 and the evaluated PTQ baselines in average accuracy across five downstream benchmarks, improving over MXFP4 by 1.24 and 0.55 on Qwen3-4B-Instruct-2507 and Qwen3-30B-A3B-Instruct-2507, respectively.
Sources
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- DeepSeek-V3 Technical Report
- MosaicQuant: Inlier-Outlier Disaggregation for Unified 4-Bit LLM Quantization
- BATQuant: Outlier-resilient MXFP4 Quantization via Learnable Block-wise Optimization
- DuQuant++: Fine-grained Rotation Enhances Microscaling FP4 Quantization
- Microscaling Data Formats for Deep Learning
- Pushing the Limits of Block Rotations in Post-Training Quantization
- Qwen3 Technical Report
- RPTQ: Reorder-based Post-training Quantization for Large Language Models
- OSC: Hardware Efficient W4A4 Quantization via Outlier Separation in Channel Dimension
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks