DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
cs.AI, cs.CL
Submitted: 2026-01-08
Updated: 2026-08-31
Comments: Accepted to COLM 2026
Code: https://github.com/google-research/mt-metrics-eval
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-Experts (MoE) has become a prominent paradigm for scaling Large Language Models (LLMs).
Terminology
Abstract
Mixture-of-Experts (MoE) has become a prominent paradigm for scaling Large Language Models (LLMs). Parameter-efficient fine-tuning methods, such as LoRA, are widely adopted to adapt pretrained MoE LLMs to downstream tasks. However, existing approaches typically assign identical LoRA ranks to all expert modules, ignoring the heterogeneous specialization of pretrained experts. This uniform allocation leads to a resource mismatch: task-relevant experts are under-provisioned, while less relevant ones receive redundant parameters. To address this, we propose DR-LoRA, a Dynamic Rank LoRA framework for fine-tuning pretrained MoE models. Specifically, DR-LoRA initializes all expert LoRA modules with a small active rank and uses an expert saliency score, which combines routing frequency and gradient-based rank importance, to identify which experts would benefit most from additional capacity. It then periodically expands the active ranks of the task-critical expert LoRA, progressively constructing a heterogeneous rank distribution tailored to the target task. Experiments on three MoE models across six tasks show that DR-LoRA consistently outperforms LoRA and other strong baselines, demonstrating that task-adaptive heterogeneous rank allocation is an effective strategy to improve active capacity utilization in MoE fine-tuning.
Sources
- Evaluating Large Language Models Trained on Code
- Training Verifiers to Solve Math Word Problems
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- LoRA+: Efficient Low Rank Adaptation of Large Models
- Mixtral of Experts
- MixLoRA: Enhancing Large Language Models Fine-Tuning with LoRA-based Mixture of Experts
- DeepSeek-V3 Technical Report
- PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model
- FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Federatedly-trained Large Language Models
- Kimi-VL Technical Report
- Qwen3 Technical Report
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
- Instruction-Following Evaluation for Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection