APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference

arXiv:2608.11688 · cs.AR, cs.AI, cs.LG · Submitted 2026-08-12 · Read on arXiv

Alish Kanani, Layan Badawi, Umit Y. Ogras

University of Wisconsin–Madison

cs.AR, cs.AI, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted at IEEE/ACM ESWEEK (CODES) 2026; the official version will appear in IEEE TCAD

License: http://creativecommons.org/licenses/by-nc-nd/4.0/

Importance score: 75/100

The gist: APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference addresses the memory bottleneck in Mixture-of-Experts (MoE) model inference at the edge.

Terminology

Summary

APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference addresses the memory bottleneck in Mixture-of-Experts (MoE) model inference at the edge. MoE models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading on the critical path. The paper notes that expert loading contributes 43% of latency and 29% of total energy for single-token generation with the Granite-3.1-3B-A800M model.

The authors argue that the solution is not static memory scaling, but adaptive resource management. They present APEX, a predictive resource management framework that overlaps expert loading with useful computation. APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model. This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques.

The key insight is that most routing misses can be avoided by fetching only a small, token-dependent number of additional experts. Therefore, APEX prefetches a top-(k + δ̂(x)) set, where δ̂(x) is dynamically selected per token using a learned confidence model. This enables APEX to fetch just enough additional experts to suppress misses without paying the energy cost of aggressive over-prefetching on every token.

APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy. In correctness-preserving mode, "the system strictly follows the original router decision. If any experts in Kr are missing from the prefetched set, they are fetched while execution begins on available experts, and the final output uses the complete original expert set. In stall-free mode, if one or more routed experts are missing, stall-free mode keeps the correctly prefetched routed experts and replaces the missing experts with the same number of highest-weight candidates from the remaining prefetched set, ranked by the original router’s softmax weights."

The prefetch router is trained via KL divergence loss to match the original router's expert softmax distribution, using only forward passes through the frozen base model. The CDF model is trained using cumulative binary cross-entropy loss to estimate the probability that a given extra-prefetch budget δ is sufficient. At runtime, given a token representation x and a target coverage τ, we select the smallest prefetch budget that satisfies: δ̂(x) = min δ ∈ 0,..., N − k pδ (x) ≥ τ.

The evaluation covers four MoE models: Granite-1B (IBM Granite-3.1-1B-A400M), Granite-3B (IBM Granite-3.1-3B-A800M), Phi-7B (Microsoft Phi-mini-MoE-7B-A2.4B), and DeepSeek-16B (DeepSeek-V2-Lite-16B-A2.4B). The hardware platform is modeled as a modern edge-class accelerator consisting of a compute chiplet tightly coupled with I/O and memory chiplets, with all expert weights reside in off-package LPDDR5X memory accessed over PCIe 6.0 x16 with 256 GB/s bidirectional bandwidth. The compute arrays provide 24 TFLOPS @ 750 MHz with 8 MB SRAM.

Key results include:

  • Across all models, APEX consistently maintains over 97% overlap across layers, compared to ProMoE which exhibits noticeable degradation in layers 5 and 8 of Granite-1B, with ∼79% and ∼84% overlap.

  • Increasing τ monotonically increases the overlap percentage across all models, with near-complete overlap (>99%) achievable with modest increases in prefetch budget.

  • APEX closely follows the ground truth when compared to an Oracle that stores the ground truth extra-prefetch budget.

  • For application-level accuracy, the impact on the perplexity is negligible for the Granite-1B, 3B and DeepSeek-16B models, especially at higher thresholds. For example, for Granite-1B moving from correctness-preserving execution to stall-free at τ = 0.90 results in only a marginal increase in perplexity (7.88 → 7.99) and a small drop in average downstream accuracy (43.3 → 42.8). However, Phi-7B exhibits a more noticeable degradation in stall-free mode, particularly at lower thresholds, because Phi models use only k = 2 experts per token, making them more sensitive to mispredictions.

  • For latency, APEX correctness-preserving mode achieves 11.41 ms average per-token latency at 512 context length for Granite-3B, which is 42% lower than no prefetching (19.77 ms) and 26% lower than ProMoE (15.39 ms).

  • For energy, APEX correctness-preserving mode consumes 287.3 mJ energy at 512 context length, which is 9.5% lower than no prefetching and 5.8% lower than ProMoE. The energy breakdown shows APEX reduces exposed stall time and lowers idle/leakage energy from 57 mJ to 8 mJ, which more than offsets the added I/O energy.

  • For EDP, APEX correctness-preserving mode lowers the EDP by 36%–42% over no prefetching, and 16%–27% over ProMoE for Granite-1B. For Granite-3B, APEX achieves 22%–30% lower EDP than ProMoE. For Phi-7B, the correctness-preserving mode achieves 48%–49% lower EDP than no prefetching, and 24%–28% lower EDP than ProMoE. For DeepSeek-16B, APEX reduces EDP by 30–41% over ProMoE.

  • The overhead is minimal: APEX adds only 0.051%, 0.046%, 0.009% and 0.036% performance overhead for Granite-1B, Granite-3B, Phi-7B, and DeepSeek-16B, respectively, with additional parameters accounting for just 0.060%, 0.027% and 0.022% of total parameters for Granite-3B, Phi-7B and DeepSeek-16B, respectively.

  • Sensitivity analysis shows APEX consistently decreases the latency by 14%–42% across the entire bandwidth range from 32 to 1024 GB/s, and APEX continues to reduce latency across all datatypes from 4-bit to 32-bit expert weights.

The paper concludes that APEX demonstrates that adaptive, confidence-aware prefetching is key to unlocking efficient MoE inference on edge systems, bridging the gap between predictive accuracy and system performance.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Resource Prefetching via Learned Confidence Models
  • AI systems can dynamically predict the minimum additional resources (e.g., memory, compute, network bandwidth) needed per input token, rather than using fixed over-provisioning. This reduces energy and latency while maintaining high accuracy.

  • Example: An edge AI serving LLMs can prefetch only the necessary expert weights for each token, cutting I/O energy by 10% and latency by 42% compared to no prefetching.

  1. Correctness-Preserving vs. Stall-Free Execution Modes
  • Systems can offer two operational modes: one that guarantees exact model semantics (for safety-critical tasks) and one that trades negligible accuracy for zero stalls (for high-throughput tasks). This enables flexible deployment across latency-sensitive and accuracy-critical applications.

  • Example: A chatbot can use correctness-preserving mode for legal/medical queries, and stall-free mode for casual conversation, with only 1.4% perplexity increase at τ=0.90.

  1. Token-Dependent Prefetch Budget Selection
  • Instead of prefetching a fixed top-k experts, AI systems can learn a per-token distribution over prefetch budgets (δ̂(x)) using a lightweight CDF model trained with cumulative binary cross-entropy. This minimizes wasted I/O while suppressing routing misses.

  • Example: For Granite-3B, this reduces exposed stall energy from 57 mJ to 8 mJ, improving energy-delay product by 22–30% over state-of-the-art ProMoE.

  1. Overlap of I/O with Computation via Predictive Routing
  • AI inference pipelines can prefetch expert weights before the attention block, using a small router that predicts the original router’s softmax distribution (trained via KL divergence). This hides memory latency behind compute, achieving >99% overlap accuracy across layers.

  • Example: On a 24 TFLOPS edge accelerator with 256 GB/s memory, this yields 11.41 ms average per-token latency for Granite-3B—42% lower than no prefetching.

  1. Robustness to Model Architecture and Hardware Variations
  • The prefetching framework generalizes across models with different expert counts (k=2 for Phi-7B, k=8 for DeepSeek-16B) and hardware bandwidths (32–1024 GB/s), consistently reducing latency by 14–42%. This enables deployment on diverse edge devices without re-tuning.

  • Example: The same APEX framework works on 4-bit to 32-bit quantized experts, maintaining performance gains across datatypes.

  1. Minimal Overhead Integration
  • The added prefetch router and CDF model introduce <0.06% parameter overhead and <0.05% compute overhead, making it feasible to integrate into existing MoE inference stacks without significant redesign.

  • Example: A 3B-parameter MoE model gains 26% latency reduction with only 0.027% extra parameters, enabling real-time edge deployment.

What the Improved AI System Can Do:

  • Run large MoE models (up to 16B parameters) on memory-constrained edge devices with near-optimal latency and energy efficiency.

  • Dynamically adapt prefetching behavior per token, balancing accuracy and resource use in real time.

  • Guarantee exact model outputs when needed, or gracefully degrade to stall-free operation with negligible quality loss.

  • Operate across diverse hardware (bandwidth, datatype, compute) without manual tuning, making edge AI more accessible and sustainable.

Abstract

Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path. We present APEX: Adaptive Expert Prefetching, a predictive resource management framework that overlaps expert loading with useful computation. APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model. This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques. APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy. Across multiple MoE models, the correctness-preserving mode reduces per-token latency by up to 26% and improves energy-delay product (EDP) by up to 41% over state-of-the-art baselines, while the stall-free mode provides additional efficiency gains with negligible impact on application accuracy. These results establish adaptive, confidence-driven expert prefetching as an effective approach for efficient MoE inference on edge systems.

Sources

Related papers