RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning
Yibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan
Ant Group
cs.DC, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-14
Code: https://github.com/areal-project/AReaL
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems.
Terminology
Summary
Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems. Sequence composition determines the dense attention work of each data-parallel (DP) microbatch, while token routing determines the sparse expert work of each expert-parallel (EP) rank. Optimizing either in isolation can shift the bottleneck to the other. In MoE RL, however, rollout-time routing replay reveals both the sequence length and the layer-wise expert demand of every sample before the corresponding training step is scheduled.
We present RoutePack, a hierarchical planner that coordinates state-consistent, layer-wise expert rerouting with attention-aware data packing over an optimizer-step window. RoutePack first uses aggregate routing demand to place experts independently at each MoE layer. It then packs all samples into the smallest certified, or best-known feasible, number of token-capped execution rows and searches their DP layout with a projected expert-data-parallel (EDP)-shard-aware objective. The objective combines a linear–quadratic attention proxy normalized over the window with the busiest physical EP rank at every MoE layer, and minimizes the accumulated cost of the slowest EDP shard. A diverse feasible population is refined through parallel population annealing without changing the selected row count, sample coverage, capacity, or communicator topology. State-consistent expert materialization preserves logical top-k routing and existing MoE kernels without introducing microbatch-level expert replication.
Across Ling-3.0-Tiny and Ling-3.0-Flash models, expert rerouting improves trainer-measured token throughput by 3.80% and 10.50%, and routing-aware packing adds another 4.86% and 3.98%, respectively. Online traces show consistent reductions in accumulated EP peaks, worst row-local peaks, and the joint bottleneck. We further derive a sufficient runtime condition under which CPU packing does not extend the training-admission critical path alongside placement-aware model-state materialization.
Improvements for AI systems
Improvements to AI systems:
-
Hierarchical joint optimization of data packing and expert routing in MoE RL training – Instead of optimizing sequence packing and expert load balancing independently, the system now coordinates both over a multi-step window, using rollout-time routing replay to pre-place experts per layer and pack samples into minimal token-capped rows. This eliminates the bottleneck-shifting problem where fixing one imbalance worsens the other.
-
State-consistent, layer-wise expert rerouting without kernel changes – The system can dynamically reassign experts across MoE layers per training step while preserving logical top-k routing and existing kernels, avoiding microbatch-level expert replication. This enables flexible load balancing without sacrificing model correctness or requiring new hardware/software stacks.
-
Projected EDP-shard-aware packing objective with linear-quadratic attention proxy – The system now predicts the combined cost of attention (via a normalized linear-quadratic proxy) and expert execution (via the busiest physical EP rank per layer), then minimizes the accumulated cost of the slowest shard. This allows the trainer to proactively avoid worst-case stragglers across both compute types.
-
Parallel population annealing for feasible packing search – The system can refine a diverse set of candidate packings without changing row count, sample coverage, token capacity, or communicator topology. This yields near-optimal or certified-feasible packings in practical time, improving throughput without sacrificing training correctness.
-
CPU-side packing with placement-aware model-state materialization – The system can offload packing computation to CPU without extending the training-admission critical path, provided a sufficient runtime condition is met. This enables asynchronous, non-blocking planning that keeps GPUs busy.
What the improved AI system can do:
-
Achieve higher token throughput in MoE RL training – Measured improvements of 3.80%–10.50% from expert rerouting alone, plus an additional 3.98%–4.86% from routing-aware packing, across models like Ling-3.0-Tiny and Ling-3.0-Flash.
-
Reduce accumulated expert-parallel peaks and worst-row-local peaks – The system consistently lowers the joint bottleneck across online traces, meaning fewer stalls and more predictable training steps.
-
Scale MoE RL to larger models without sacrificing load balance – By jointly planning over a window, the system handles the dynamic, non-stationary routing demands typical of RL rollouts, which static packing or routing methods fail to address.
-
Maintain model fidelity – State-consistent materialization ensures that top-k routing semantics are preserved, so the improved throughput does not come at the cost of altered model behavior or degraded policy learning.
-
Operate within existing infrastructure – No changes to communicator topology, expert kernels, or microbatch-level replication are required, making the improvement directly deployable to current MoE RL training stacks.
Sources
- GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- ReLibra: Routing-Replay-Guided Load Balancing for MoE Training in Reinforcement Learning
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Semantic Parallelism: Redefining Efficient MoE Inference via Model-Data Co-Scheduling
- Stabilizing MoE Reinforcement Learning by Aligning Training and Inference Routers
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Coordinated Scheduling for MoE LLM Serving
- Libra: Taming Attention Workload Skew in Long-Context LLM Training with Bounded Sequence Pool
- UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing
- mHC: Manifold-Constrained Hyper-Connections
- Fine-grained MoE Load Balancing with Linear Programming
- Harnessing Routing Foresight for Micro-step-level MoE load balancing in RL Post-training
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing