Preserving Long-Tailed Expert Information in Mixture-of-Experts Tuning
cs.LG, cs.CL
Submitted: 2026-04-24
Updated: 2026-09-08
Comments: Camera-ready version. Accepted at the Third Conference on Language Modeling (COLM 2026)
Journal ref: Proceedings of the Third Conference on Language Modeling (COLM 2026), 2026
Code: https://github.com/HectorHHZ/ExpertCondenser
License: http://creativecommons.org/licenses/by/4.0/
The gist: Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layers are fragile.
Terminology
Abstract
Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layers are fragile. Methods such as DenseMixer and ESFT mitigate router collapse with dense mixing or auxiliary load-balancing losses, but these introduce noisy gradients that often degrade performance. In preliminary experiments, we systematically pruned experts and observed that while certain super experts are activated far more frequently, discarding less used experts still leads to notable performance degradation. This suggests that even rarely activated experts encode non-trivial knowledge useful for downstream tasks. Motivated by this, we propose an auxiliary-loss-free MoE SFT framework that combines bias-driven sparsification with always-active gated condenser experts. Rather than enforcing balanced activation across all experts, our method encourages task-relevant experts to remain active while pushing long-tailed experts toward inactivity. The condenser experts provide a persistent, learnable pathway that alleviates gradient starvation and facilitates consolidation of information that would otherwise remain fragmented across sparsely activated experts. Analysis further suggest that this design better preserves long-tailed expert information under sparse routing. Experiments on large-scale MoE models demonstrate that our approach outperforms state-of-the-art SFT baselines such as DenseMixer and ESFT, achieving average gain of 2.5%+ on both mathematical reasoning and commonsenseQA benchmarks.
Sources
- Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
- Training Verifiers to Solve Math Word Problems
- Sparse Matrix in Large Language Model Fine-tuning
- LoRA: Low-Rank Adaptation of Large Language Models
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language Models
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
- DeepSeek-V3 Technical Report
- Improving Generalization in Aerial and Terrestrial Mobile Robots Control Through Delayed Policy Learning
- DoRA: Weight-Decomposed Low-Rank Adaptation
- Not All Experts are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models
- OLMoE: Open Mixture-of-Experts Language Models
- s1: Simple test-time scaling
- ZeRO-Offload: Democratizing Billion-Scale Model Training
- Solving General Arithmetic Word Problems
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
- Let the Expert Stick to His Last: Expert-Specialized Fine-Tuning for Sparse Architectural Large Language Models
- mHC: Manifold-Constrained Hyper-Connections
- Qwen2 Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks