How to Loop MoE: Flatten the Experts, Untie the Attention
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/SR-A-W/how-to-loop-moe
Terminology
Sources
- $\phi$-Balancing for Mixture-of-Experts Training
- LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
- Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
- DeepSeek-V3 Technical Report
- AI and Memory Wall
- GLM-5: from Vibe Coding to Agentic Engineering
- Mixture of A Million Experts
- Tying the Loop -- Tied Expert Layers in Mixture-of-Experts Language Models
- Mixtral of Experts
- Kimi K2: Open Agentic Intelligence
- Kimi K3: Open Frontier Intelligence
- Megrez2 Technical Report
- The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
- Parcae: Scaling Laws For Stable Looped Language Models
- MoRE: Mixture of Reused Experts
- Qwen3 Technical Report
- How Much Is One Recurrence Worth? Iso-Depth Scaling Laws for Looped Language Models
- From Score Distributions to Balance: Plug-and-Play Mixture-of-Experts Routing
- ReXMoE: Reusing Experts with Minimal Overhead in Mixture-of-Experts
- Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks