From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs
cs.LG, cs.AI
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/aheadformore/NSFT
License: http://creativecommons.org/licenses/by/4.0/
The gist: As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models.
Terminology
Abstract
As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question for parameter-efficient fine-tuning (PEFT): at what granularity should parameters be selected and updated? Existing PEFT methods such as LoRA operate on predefined weight matrices, while expert-level sparse tuning methods update entire selected experts. However, we observe that activated experts are internally sparse, with only a small fraction of intermediate channels strongly responding to downstream tasks, indicating that expert-level adaptation is still too coarse. We propose NSFT (Neural Sub-expert Fine-Tuning), a fine-grained PEFT framework that refines MoE adaptation from experts to sub-experts. NSFT decomposes each expert along the intermediate dimension into structured channel groups and selects task-relevant sub-experts by combining routing importance with intra-expert activation saliency. To optimize sparse partial updates, NSFT further introduces learning-rate scaling and dynamic gradient scaling to compensate for the reduced effective update magnitude. Experiments on OLMoE and Ling-mini-2.0 across challenging domain-specific tasks and general benchmarks show that NSFT consistently outperforms representative PEFT and expert-level sparse tuning baselines, while using substantially fewer trainable parameters and preserving competitive general capability. These results suggest that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.
Sources
- Program Synthesis with Large Language Models
- Qwen Technical Report
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- Evaluating Large Language Models Trained on Code
- Mixture of Neuron Experts
- Training Verifiers to Solve Math Word Problems
- DR-LoRA: Dynamic Rank LoRA for Fine-Tuning Mixture-of-Experts Models
- The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
- The Llama 3 Herd of Models
- Measuring Mathematical Problem Solving With the MATH Dataset
- Mixtral of Experts
- DeepSeek-V3 Technical Report
- OLMoE: Open Mixture-of-Experts Language Models
- Olmo 3
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Kimi K2: Open Agentic Intelligence
- Not All Correct Answers Are Equal: Why Your Distillation Source Matters
- AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning
- Qwen3 Technical Report
- Instruction-Following Evaluation for Large Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks