Ban&Pick: Enhancing Performance and Efficiency of MoE-LLMs via Smarter Routing
cs.LG, cs.AI
Submitted: 2025-09-08
Updated: 2026-08-26
Comments: 26 pages, 13 figures
Journal ref: The 2026 Conference on Empirical Methods in Natural Language Processing (EMNLP 2026)
Code: https://github.com/open-compass/opencompass
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Toward Inference-optimal Mixture-of-Expert Large Language Models
- Mixture Compressor for Mixture-of-Experts LLMs Gains More
- Mixtral of Experts
- Let's Verify Step by Step
- Training Verifiers to Solve Math Word Problems
- Program Synthesis with Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks