MoRE: Scaling mixture of experts with hardware-aware low-rank routing
cs.LG, cs.AI, cs.CL, stat.ML
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/Matheart/MoRE_code
Terminology
Sources
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language Models
- Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
- Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws
- Memory Layers at Scale
- Secret mixtures of experts inside your LLM
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- On the Representation Collapse of Sparse Mixture of Experts
- Unified Scaling Laws for Routed Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- SonicMoE: Accelerating MoE with IO and Tile-aware Optimizations
- LoRA+: Efficient Low Rank Adaptation of Large Models
- ExpertFlow: Efficient Mixture-of-Experts Inference via Predictive Expert Caching and Token Scheduling
- Mixture of A Million Experts
- Liger Kernel: Efficient Triton Kernels for LLM Training
- LoRA: Low-Rank Adaptation of Large Language Models
- Ultra-Sparse Memory Network
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks