Structure Aware Neural Architecture Search for Mixture of Experts
cs.LG
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/intsystems/nas-for-moe
Terminology
Sources
- A Comprehensive Survey of Mixture-of-Experts: Algorithms, Theory, and Applications
- Mixtral of Experts
- DARTS: Differentiable Architecture Search
- Mixture of experts models for multilevel data: modelling framework and approximation theory
- Multi-Source Domain Adaptation with Mixture of Experts
- A General Theory for Softmax Gating Multinomial Logistic Mixture of Experts
- Divide, Specialize, and Route: A New Approach to Efficient Ensemble Learning
- Learning Factored Representations in a Deep Mixture of Experts
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Scaling Vision with Sparse Mixture of Experts
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- Demons in the Detail: On Implementing Load Balancing Loss for Training Specialized Mixture-of-Expert Models
- Branch-Train-Merge: Embarrassingly Parallel Training of Expert Language Models
- Scaling Expert Language Models with Unsupervised Domain Discovery
- Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM
- OLMoE: Open Mixture-of-Experts Language Models
- Time-MoE: Billion-Scale Time Series Foundation Models with Mixture of Experts
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
- Neural Architecture Search with Reinforcement Learning
- Designing Neural Network Architectures using Reinforcement Learning
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks