Function Over Form: Distributional Orthogonalization in Mixture-of-Experts with Replica Expert Mechanism
cs.AI
Submitted: 2026-09-26
Updated: 2026-09-26
Terminology
Sources
- Prediction Is All MoE Needs: Expert Load Distribution Goes from Fluctuating to Stabilizing
- Scalable and Flexible Causal Discovery with an Efficient Test for Adjacency
- Mixtral of Experts
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Diversifying the Mixture-of-Experts Representation for Language Models with Orthogonal Optimizer
- DeepSeek-V3 Technical Report
- Texture-guided Coding for Deep Features
- Coupling Experts and Routers in Mixture-of-Experts via an Auxiliary Loss
- OLMoE: Open Mixture-of-Experts Language Models
- GLU Variants Improve Transformer
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
- Qwen3 Technical Report
- ST-MoE: Designing Stable and Transferable Sparse Expert Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection