Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts

arXiv:2608.10605 · cs.LG, cs.AI · Submitted 2026-08-11 · Read on arXiv

Soumajyoti Sarkar, Yuxin Tang, Sheng Zha

Amazon AGI Foundations

cs.LG, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/dmlc/ScalePlan

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: This paper introduces MOSAIC (Model Optimization via Systems-Aware TraIning Co-design), a framework that jointly optimizes model architecture, training-token budget, and distributed parallel

Terminology

Summary

This paper introduces MOSAIC (Model Optimization via Systems-Aware TraIning Co-design), a framework that jointly optimizes model architecture, training-token budget, and distributed parallel execution layout for a fixed cluster and training window, specifically instantiated for sparse Mixture-of-Experts (MoE) language models. The central thesis is that conventional compute-optimal scaling laws, which optimize loss under a model-FLOPs budget, are insufficient for sparse MoEs because they ignore systems efficiency. The paper argues for a shift from a model-FLOPs budget to a hardware-deliverable model-FLOPs budget.

The paper makes the following key contributions and observations:

  1. New Scaling Law with Systems-Side Knobs: The authors fit a new four-dimensional joint MoE pretraining scaling law over total parameters (Ntot), sparsity (S), training tokens (D), and expert split factor (G), defined as:

L(Ntot, S, D, G) = a/Ntot α + b/D β + c/(1-S) λ + j/((1-S) δ * Ntot γ * G η) + e.

This law is fit on 150 from-scratch sparse MoE pretraining runs spanning active parameters from 104 million to 2.7 billion and total model sizes up to 79 billion parameters. The expert split factor G is identified as a necessary axis, with the fitted exponent η ≈ +0.95, indicating that finer-grained experts lower loss at fixed capacity.

  1. Boundary-Seeking Model-FLOPs Optimum: Through scaling analysis, the paper shows that within the calibrated sparsity range (S ∈ [0.5, 0.981]), a model-FLOPs budget admits no interior optimal sparsity. The fitted loss decreases monotonically with sparser models, so the compute optimum lies at the upper boundary of the data support (S = Smax). This is because larger sparsity increases optimal total parameters, decreases optimal active parameters, and lowers achievable loss at every model-FLOPs budget.

  2. Operator-Level Performance Model: The authors develop an analytical performance model (released as ScalePlan) that predicts MFU, memory footprint, and the best parallel layout. It uses microbenchmarks for GEMMs, grouped GEMMs (expert MLPs), attention, routing, cross-entropy, and NCCL collectives, combined with an iteration-time model that accounts for compute vs. dispatch, pipeline bubbles, and vocabulary-stage imbalance. The model is validated against training runs up to 18 billion active parameters, with a mean absolute percentage error of predicted MFU under 15% per sweep.

  3. MOSAIC Optimization and Results: MOSAIC couples the scaling law and performance model to search the architecture grid on a cluster of NVIDIA B200 GPUs under a fixed GPU-hour budget. Key results include:

  • The lowest-loss model configuration is not the one that emits the most model FLOPs. Under a 32-node, 20-day envelope, the optimal configuration is a q=12, E=96, K=2, G=4 model with Nact=14.5B, S=0.956, running at 12.4% predicted MFU and loss L=1.3898. The candidate that realizes the most model FLOPs (a denser G=2, E=32 geometry) lands at a higher loss.

  • The optimal sparsity S* becomes interior and hardware-dependent (roughly 0.915–0.963) once systems constraints are added, rather than following the monotone boundary preference implied by model FLOPs.

  • Staged pretraining runs of up to 250B total parameters validate the ranked selections. The predicted MFU ranking reproduces on hardware, and the loss ordering flips from the model-FLOPs axis to the peak-equivalent hardware-compute axis, as MOSAIC predicts.

The paper concludes that the separation of scaling law and systems tuning is consequential at scale, and argues for unified architecture and systems co-design for frontier language model training. Limitations include the dependence on the chosen geometry ladder, hardware platform, and training recipe, and the substantial extrapolation required for some reported optima.

Improvements for AI systems

Based on this paper, here are the specific improvements I can make to AI systems, and what the improved systems can do:

1. Hardware-Aware Architecture Search for MoE Models

I can now automatically select the optimal sparsity, number of experts, expert split factor, and active parameter count for a given cluster and training deadline, rather than relying on generic FLOPs-based scaling laws. The improved system will train to a lower loss within the same GPU-hour budget by explicitly modeling memory bandwidth, dispatch overhead, and pipeline bubbles.

2. Joint Optimization of Training Schedule and Parallel Layout

I can co-optimize the token budget, expert parallelism degree, tensor parallelism, and pipeline stages simultaneously. The improved system will avoid configurations that maximize raw model FLOPs but underutilize hardware (e.g., the paper shows a 12.4% MFU config beats a denser one with higher FLOPs), leading to faster wall-clock convergence for a fixed cluster.

3. Boundary-Aware Sparsity Selection

I can now predict that under a pure compute budget, sparsity should be pushed to the maximum supported value (e.g., S=0.981), but under real hardware constraints, the optimal sparsity becomes interior (0.915–0.963). The improved system will automatically detect this shift and select a sparsity that balances expert granularity gains against dispatch and memory costs, avoiding over-sparsification that hurts MFU.

4. Fine-Grained Expert Splitting for Lower Loss

Using the fitted exponent η≈0.95, I can now decide to increase the expert split factor G (e.g., from 2 to 4) even if it increases communication overhead, because the loss reduction from finer-grained experts outweighs the systems cost in many regimes. The improved system will generate MoE configurations with more, smaller experts that achieve better loss-per-parameter than coarser alternatives.

5. Predictive Performance Modeling for Rapid Iteration

I can use the ScalePlan operator-level model to predict MFU, memory, and optimal parallel layout for any candidate MoE architecture before training, with <15% error. The improved system will run thousands of virtual architecture sweeps in seconds, filtering out infeasible or inefficient designs without needing expensive trial runs, and then validate only the top-ranked candidates on real hardware.

6. Staged Pretraining with Ranked Configuration Selection

I can now run a two-stage process: first, use the joint scaling law to rank candidate configurations by predicted loss under a hardware-deliverable FLOPs budget; second, validate the top few on a small-scale run (e.g., 250B parameters) to confirm the ranking. The improved system will reliably pick the best architecture for a large-scale run, avoiding the common failure of choosing a model that looks good on paper but trains poorly in practice.

7. Cluster-Specific Optimization

I can tailor the architecture and parallel layout to the exact interconnect bandwidth, GPU memory, and compute capability of the target cluster (e.g., NVIDIA B200s). The improved system will automatically adapt its recommendations if the cluster changes, rather than assuming a universal optimal model size or sparsity.

Abstract

In large-scale pretraining, the algorithm, architecture, and systems decisions are conventionally made in disconnected stages. A scaling law stage selects an architecture and training recipe, optimizing loss under compute constraints, and a separate systems stage then optimizes the implementation for hardware efficiency. In this work, we develop MOSAIC, which formulates model architecture and systems co-design as an optimization problem. MOSAIC couples a predictive scaling law with a calibrated performance model that estimates Model FLOPs Utilization (MFU), communication cost, memory footprint, and the best parallel layout. We instantiate the framework for sparse Mixture-of-Experts (MoE) language models, where expert count, routing sparsity, and other MoE layer dimensions affect both the loss and systems efficiency. We fit a scaling law on sparse MoE models trained on text data, whose scaling dimensions include the sparsity factor, which is the fraction of model parameters inactive per token in a forward pass. The scaling law sweeps in our work span active parameters from 104 million to 2.7 billion and total model sizes reaching 79 billion parameters. We show that, within the calibrated sparsity range, an efficiency-agnostic model-FLOPs budget admits no interior optimal sparsity. The fitted loss decreases monotonically with sparser models and the compute optimum lies at the upper boundary of the data support. An optimal sparsity in MoE models instead emerges under the cluster's systems constraints, as captured by MOSAIC. Our results argue for a shift towards unified architecture and systems co-design for frontier language model training.

Sources

Related papers