RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning

arXiv:2608.12146 · cs.DC, cs.LG · Submitted 2026-08-12 · Read on arXiv

Yibo Shen, Xudong Han, Xiaowei Zhu, Gen Li, Zhenxuan Pan

Ant Group

cs.DC, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-14

Code: https://github.com/areal-project/AReaL

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems.

Terminology

Summary

Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems. Sequence composition determines the dense attention work of each data-parallel (DP) microbatch, while token routing determines the sparse expert work of each expert-parallel (EP) rank. Optimizing either in isolation can shift the bottleneck to the other. In MoE RL, however, rollout-time routing replay reveals both the sequence length and the layer-wise expert demand of every sample before the corresponding training step is scheduled.

We present RoutePack, a hierarchical planner that coordinates state-consistent, layer-wise expert rerouting with attention-aware data packing over an optimizer-step window. RoutePack first uses aggregate routing demand to place experts independently at each MoE layer. It then packs all samples into the smallest certified, or best-known feasible, number of token-capped execution rows and searches their DP layout with a projected expert-data-parallel (EDP)-shard-aware objective. The objective combines a linear–quadratic attention proxy normalized over the window with the busiest physical EP rank at every MoE layer, and minimizes the accumulated cost of the slowest EDP shard. A diverse feasible population is refined through parallel population annealing without changing the selected row count, sample coverage, capacity, or communicator topology. State-consistent expert materialization preserves logical top-k routing and existing MoE kernels without introducing microbatch-level expert replication.

Across Ling-3.0-Tiny and Ling-3.0-Flash models, expert rerouting improves trainer-measured token throughput by 3.80% and 10.50%, and routing-aware packing adds another 4.86% and 3.98%, respectively. Online traces show consistent reductions in accumulated EP peaks, worst row-local peaks, and the joint bottleneck. We further derive a sufficient runtime condition under which CPU packing does not extend the training-admission critical path alongside placement-aware model-state materialization.

Improvements for AI systems

Improvements to AI systems:

  1. Hierarchical joint optimization of data packing and expert routing in MoE RL training – Instead of optimizing sequence packing and expert load balancing independently, the system now coordinates both over a multi-step window, using rollout-time routing replay to pre-place experts per layer and pack samples into minimal token-capped rows. This eliminates the bottleneck-shifting problem where fixing one imbalance worsens the other.

  2. State-consistent, layer-wise expert rerouting without kernel changes – The system can dynamically reassign experts across MoE layers per training step while preserving logical top-k routing and existing kernels, avoiding microbatch-level expert replication. This enables flexible load balancing without sacrificing model correctness or requiring new hardware/software stacks.

  3. Projected EDP-shard-aware packing objective with linear-quadratic attention proxy – The system now predicts the combined cost of attention (via a normalized linear-quadratic proxy) and expert execution (via the busiest physical EP rank per layer), then minimizes the accumulated cost of the slowest shard. This allows the trainer to proactively avoid worst-case stragglers across both compute types.

  4. Parallel population annealing for feasible packing search – The system can refine a diverse set of candidate packings without changing row count, sample coverage, token capacity, or communicator topology. This yields near-optimal or certified-feasible packings in practical time, improving throughput without sacrificing training correctness.

  5. CPU-side packing with placement-aware model-state materialization – The system can offload packing computation to CPU without extending the training-admission critical path, provided a sufficient runtime condition is met. This enables asynchronous, non-blocking planning that keeps GPUs busy.

What the improved AI system can do:

  • Achieve higher token throughput in MoE RL training – Measured improvements of 3.80%–10.50% from expert rerouting alone, plus an additional 3.98%–4.86% from routing-aware packing, across models like Ling-3.0-Tiny and Ling-3.0-Flash.

  • Reduce accumulated expert-parallel peaks and worst-row-local peaks – The system consistently lowers the joint bottleneck across online traces, meaning fewer stalls and more predictable training steps.

  • Scale MoE RL to larger models without sacrificing load balance – By jointly planning over a window, the system handles the dynamic, non-stationary routing demands typical of RL rollouts, which static packing or routing methods fail to address.

  • Maintain model fidelity – State-consistent materialization ensures that top-k routing semantics are preserved, so the improved throughput does not come at the cost of altered model behavior or degraded policy learning.

  • Operate within existing infrastructure – No changes to communicator topology, expert kernels, or microbatch-level replication are required, making the improvement directly deployable to current MoE RL training stacks.

Sources

Related papers