Scaling Laws for Looped Mixture of Experts
cs.LG, cs.AI, cs.CL
Submitted: 2026-09-30
Updated: 2026-09-30
Terminology
Sources
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
- Relaxed Recursive Transformers: Effective Parameter Sharing with Layer-wise LoRA
- MobileMoE: Scaling On-Device Mixture of Experts
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V3 Technical Report
- Universal Transformers
- Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- Looped Transformers as Programmable Computers
- ELT: Elastic Looped Transformers for Visual Generation
- Measuring Massive Multitask Language Understanding
- Training Compute-Optimal Large Language Models
- LoopFormer: Elastic-Depth Looped Transformers for Latent Reasoning via Shortcut Modulation
- Mixtral of Experts
- Scaling Laws for Neural Language Models
- Scaling Laws for Fine-Grained Mixture of Experts
- Sparse Layers are Critical to Scaling Looped Language Models
- Joint MoE Scaling Laws: Mixture of Experts Can Be Memory Efficient
- Teaching Pretrained Language Models to Think Deeper with Retrofitted Recurrence
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks