MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs
cs.LG, cs.AI
Submitted: 2026-06-15
Updated: 2026-08-26
Comments: 19 pages, 8 figures
Journal ref: The 2026 Conference on Empirical Methods in Natural Language Processing, Main Conference (EMNLP 2026 Main)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
- LLaVA-OneVision: Easy Visual Task Transfer
- MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
- MBQ: Modality-Balanced Quantization for Large Vision-Language Models
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Evaluating Object Hallucination in Large Vision-Language Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- MoE-LLaVA: Mixture of Experts for Large Vision-Language Models
- SPEED-Q: Staged Processing with Enhanced Distillation towards Efficient Low-bit On-device VLM Quantization
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
- Mixture Compressor for Mixture-of-Experts LLMs Gains More
- InfographicVQA
- VEQ: Modality-Adaptive Quantization for MoE Vision-Language Models
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
- Kimi-VL Technical Report
- DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset Adaptation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks