Dynamic Expert Quantization for Scalable Mixture-of-Experts Inference
cs.PF, cs.AI, cs.LG
Submitted: 2025-11-19
Updated: 2026-09-10
Comments: 13 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
- MxMoE: Mixed-precision Quantization for MoE with Accuracy and Performance Co-Design
- Fast Inference of Mixture-of-Experts Language Models with Offloading
- QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization
- MoEQuant: Enhancing Quantization for Mixture-of-Experts Large Language Models via Expert-Balanced Sampling and Affinity Guidance
- Mixture of Quantized Experts (MoQE): Complementary Effect of Low-bit Quantization and Robustness
- FlexRound: Learnable Rounding based on Element-wise Division for Post-Training Quantization
- Adaptive Gating in Mixture-of-Experts based Language Models
- Pointer Sentinel Mixture Models
- ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
- ProMoE: Fast MoE-based LLM Serving using Proactive Caching
- HOBBIT: A Mixed Precision Expert Offloading System for Fast MoE Inference
- Qwen3 Technical Report
- HAQ: Hardware-Aware Automated Quantization with Mixed Precision
- MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache