LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling
cs.LG, cs.AI
Submitted: 2026-06-03
Updated: 2026-08-27
Journal ref: EMNLP 2026
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth.
Terminology
Abstract
Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth. However, mainstream looped architectures rely on dense backbones that couple parameter count with per-token FLOPs, which makes it impossible to isolate the effect of iterative computation under matched budgets. To this end, we present LoopMoE, a looped MoE language model that integrates sparse routing with iterative weight-shared computation through two designs. The first is IterAdaLN, which resolves weight-sharing symmetry via a modulation signal jointly conditioned on the iteration index and the per-token hidden state. The second is a capacity-balancing strategy that recovers the attention-to-FFN active parameter ratio of well-tuned non-looped references. Together, these designs enable the first strictly controlled, head-to-head evaluation of a looped MoE against a Vanilla MoE under identical total parameters, per-token FLOPs, and active sublayer ratios. Across nine downstream benchmarks, LoopMoE's average improvement over its matched vanilla MoE increases from over 1 point at the 3B scale to approximately 3 points at the 9B scale. These results provide initial evidence that the benefits of iterative sparse computation may strengthen with scale, positioning LoopMoE as a promising architecture for scalable looped language models.
Sources
- Training Verifiers to Solve Math Word Problems
- Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- Universal Transformers
- A Mechanistic Analysis of Looped Reasoning Language Models
- Training Large Language Models to Reason in a Continuous Latent Space
- Dr.LLM: Dynamic Layer Routing in LLMs
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- DeepSeek-V3 Technical Report
- Decoupled Weight Decay Regularization
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
- Mixtral of Experts
- Olmo 3
- 2 OLMo 2 Furious
- Power Scheduler: A Batch Size and Token Number Agnostic Learning Rate Scheduler
- OpenMoE: An Early Effort on Open Mixture-of-Experts Language Models
- Qwen3 Technical Report
- Hyperloop Transformers
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks