Activation Outliers Matter: Robust Recovery for Quantized Multimodal LLMs
cs.LG
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: 14 Pages, 5 figures, 5 tables
Code: https://github.com/NVIDIA/TransformerEngine
License: http://creativecommons.org/publicdomain/zero/1.0/
The gist: Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs).
Terminology
Abstract
Low-bit quantization offers a promising avenue for reducing the computational and memory demands of Multimodal Large Language Models (MLLMs). Recent hardware support for low-precision formats, ranging from MXFP8 to ultra-low-bit formats such as MXFP4 and HiF4, has accelerated research into efficient MLLM training and deployment. In this work, we present a systematic study of these quantization schemes in representative MLLMs that span both video generation and reasoning tasks. Our analysis shows that MXFP8 achieves near-lossless performance, whereas aggressive 4-bit quantization leads to significant degradation. Through extensive ablations, we identify activation quantization as the primary source of this performance loss, contributing substantially more than weight quantization. Motivated by this observation, we propose Residual Fallback Quantization (RFQ), a lightweight activation reconstruction framework that supplements the primary ulta-low-bit activation representation with an auxiliary quantized residual pathway. By explicitly modeling and compensating for quantization errors, RFQ improves activation fidelity while preserving the efficiency advantages of ultra-low-bit computation. RFQ requires no architectural modifications and incurs negligible computational overhead. Extensive experiments on Wan2.2 and Qwen3-VL demonstrate that RFQ consistently recovers a substantial portion of the performance lost under the quantization of MXFP4 and HiF4, significantly narrowing the gap to BF16 baselines across both generation and 4 reasoning benchmarks. Our findings establish activation quantization as the dominant bottleneck in ultra-low-bit MLLMs and highlight residual-based activation reconstruction as an effective and practical strategy for robust 4-bit deployment.
Sources
- Qwen3-VL Technical Report
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- BASE-Q: Bias and Asymmetric Scaling Enhanced Rotational Quantization for Large Language Models
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- HiFloat4 Format for Language Model Inference
- BitNet b1.58 2B4T Technical Report
- OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
- QuEST: Stable Training of LLMs with 1-Bit Weights and Activations
- Movie Gen: A Cast of Media Foundation Models
- Microscaling Data Formats for Deep Learning
- HiFloat4 Format for Language Model Pre-training on Ascend NPUs
- MF-QAT: Multi-Format Quantization-Aware Training for Elastic Inference
- SalQ-VLM: Fine-Grained Saliency-Guided Quantization for Vision-Language Models
- Attn-QAT: 4-Bit Attention With Quantization-Aware Training
- Accurate INT8 Training Through Dynamic Block-Level Fallback
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks