APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference
Alish Kanani, Layan Badawi, Umit Y. Ogras
University of Wisconsin–Madison
cs.AR, cs.AI, cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: Accepted at IEEE/ACM ESWEEK (CODES) 2026; the official version will appear in IEEE TCAD
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 75/100
The gist: APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference addresses the memory bottleneck in Mixture-of-Experts (MoE) model inference at the edge.
Terminology
Summary
APEX: Adaptive Expert Prefetching for Memory-Efficient Edge MoE Inference addresses the memory bottleneck in Mixture-of-Experts (MoE) model inference at the edge. MoE models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency.
However, MoE inference at the edge is fundamentally limited by memory.
Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading on the critical path. The paper notes that expert loading contributes 43% of latency and 29% of total energy
for single-token generation with the Granite-3.1-3B-A800M model.
The authors argue that the solution is not static memory scaling, but adaptive resource management.
They present APEX, a predictive resource management framework that overlaps expert loading with useful computation.
APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model.
This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques.
The key insight is that most routing misses can be avoided by fetching only a small, token-dependent number of additional experts.
Therefore, APEX prefetches a top-(k + δ̂(x)) set, where δ̂(x) is dynamically selected per token using a learned confidence model.
This enables APEX to fetch just enough additional experts to suppress misses without paying the energy cost of aggressive over-prefetching on every token.
APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy.
In correctness-preserving mode, "the system strictly follows the original router decision. If any experts in Kr are missing from the prefetched set, they are fetched while execution begins on available experts, and the final output uses the complete original expert set. In stall-free mode,
if one or more routed experts are missing, stall-free mode keeps the correctly prefetched routed experts and replaces the missing experts with the same number of highest-weight candidates from the remaining prefetched set, ranked by the original router’s softmax weights."
The prefetch router is trained via KL divergence loss
to match the original router's expert softmax distribution, using only forward passes through the frozen base model. The CDF model is trained using cumulative binary cross-entropy loss
to estimate the probability that a given extra-prefetch budget δ is sufficient. At runtime, given a token representation x and a target coverage τ, we select the smallest prefetch budget that satisfies: δ̂(x) = min δ ∈ 0,..., N − k pδ (x) ≥ τ.
The evaluation covers four MoE models: Granite-1B (IBM Granite-3.1-1B-A400M), Granite-3B (IBM Granite-3.1-3B-A800M), Phi-7B (Microsoft Phi-mini-MoE-7B-A2.4B), and DeepSeek-16B (DeepSeek-V2-Lite-16B-A2.4B). The hardware platform is modeled as a modern edge-class accelerator consisting of a compute chiplet tightly coupled with I/O and memory chiplets,
with all expert weights reside in off-package LPDDR5X memory
accessed over PCIe 6.0 x16
with 256 GB/s bidirectional bandwidth. The compute arrays provide 24 TFLOPS @ 750 MHz
with 8 MB
SRAM.
Key results include:
-
Across all models, APEX consistently maintains over 97% overlap across layers,
compared to ProMoE whichexhibits noticeable degradation in layers 5 and 8 of Granite-1B, with ∼79% and ∼84% overlap.
-
Increasing τ monotonically increases the overlap percentage across all models,
with near-complete overlap (>99%) achievable with modest increases in prefetch budget. -
APEX closely follows the ground truth
when compared to an Oracle that stores the ground truth extra-prefetch budget. -
For application-level accuracy,
the impact on the perplexity is negligible for the Granite-1B, 3B and DeepSeek-16B models, especially at higher thresholds.
For example,for Granite-1B moving from correctness-preserving execution to stall-free at τ = 0.90 results in only a marginal increase in perplexity (7.88 → 7.99) and a small drop in average downstream accuracy (43.3 → 42.8).
However,Phi-7B exhibits a more noticeable degradation in stall-free mode, particularly at lower thresholds,
becausePhi models use only k = 2 experts per token,
making themmore sensitive to mispredictions.
-
For latency,
APEX correctness-preserving mode achieves 11.41 ms average per-token latency
at 512 context length for Granite-3B, which is42% lower than no prefetching (19.77 ms) and 26% lower than ProMoE (15.39 ms).
-
For energy,
APEX correctness-preserving mode consumes 287.3 mJ energy
at 512 context length, which is9.5% lower than no prefetching and 5.8% lower than ProMoE.
The energy breakdown shows APEXreduces exposed stall time and lowers idle/leakage energy from 57 mJ to 8 mJ,
whichmore than offsets the added I/O energy.
-
For EDP,
APEX correctness-preserving mode lowers the EDP by 36%–42% over no prefetching, and 16%–27% over ProMoE
for Granite-1B. For Granite-3B, APEX achieves22%–30% lower EDP than ProMoE.
For Phi-7B,the correctness-preserving mode achieves 48%–49% lower EDP than no prefetching, and 24%–28% lower EDP than ProMoE.
For DeepSeek-16B,APEX reduces EDP by 30–41% over ProMoE.
-
The overhead is minimal:
APEX adds only 0.051%, 0.046%, 0.009% and 0.036% performance overhead for Granite-1B, Granite-3B, Phi-7B, and DeepSeek-16B, respectively,
with additional parameters accounting forjust 0.060%, 0.027% and 0.022% of total parameters for Granite-3B, Phi-7B and DeepSeek-16B, respectively.
-
Sensitivity analysis shows
APEX consistently decreases the latency by 14%–42% across the entire bandwidth range
from 32 to 1024 GB/s, andAPEX continues to reduce latency across all datatypes
from 4-bit to 32-bit expert weights.
The paper concludes that APEX demonstrates that adaptive, confidence-aware prefetching is key to unlocking efficient MoE inference on edge systems, bridging the gap between predictive accuracy and system performance.
Improvements for AI systems
Improvements to AI Systems:
- Adaptive Resource Prefetching via Learned Confidence Models
-
AI systems can dynamically predict the minimum additional resources (e.g., memory, compute, network bandwidth) needed per input token, rather than using fixed over-provisioning. This reduces energy and latency while maintaining high accuracy.
-
Example: An edge AI serving LLMs can prefetch only the necessary expert weights for each token, cutting I/O energy by 10% and latency by 42% compared to no prefetching.
- Correctness-Preserving vs. Stall-Free Execution Modes
-
Systems can offer two operational modes: one that guarantees exact model semantics (for safety-critical tasks) and one that trades negligible accuracy for zero stalls (for high-throughput tasks). This enables flexible deployment across latency-sensitive and accuracy-critical applications.
-
Example: A chatbot can use correctness-preserving mode for legal/medical queries, and stall-free mode for casual conversation, with only 1.4% perplexity increase at τ=0.90.
- Token-Dependent Prefetch Budget Selection
-
Instead of prefetching a fixed top-k experts, AI systems can learn a per-token distribution over prefetch budgets (δ̂(x)) using a lightweight CDF model trained with cumulative binary cross-entropy. This minimizes wasted I/O while suppressing routing misses.
-
Example: For Granite-3B, this reduces exposed stall energy from 57 mJ to 8 mJ, improving energy-delay product by 22–30% over state-of-the-art ProMoE.
- Overlap of I/O with Computation via Predictive Routing
-
AI inference pipelines can prefetch expert weights before the attention block, using a small router that predicts the original router’s softmax distribution (trained via KL divergence). This hides memory latency behind compute, achieving >99% overlap accuracy across layers.
-
Example: On a 24 TFLOPS edge accelerator with 256 GB/s memory, this yields 11.41 ms average per-token latency for Granite-3B—42% lower than no prefetching.
- Robustness to Model Architecture and Hardware Variations
-
The prefetching framework generalizes across models with different expert counts (k=2 for Phi-7B, k=8 for DeepSeek-16B) and hardware bandwidths (32–1024 GB/s), consistently reducing latency by 14–42%. This enables deployment on diverse edge devices without re-tuning.
-
Example: The same APEX framework works on 4-bit to 32-bit quantized experts, maintaining performance gains across datatypes.
- Minimal Overhead Integration
-
The added prefetch router and CDF model introduce <0.06% parameter overhead and <0.05% compute overhead, making it feasible to integrate into existing MoE inference stacks without significant redesign.
-
Example: A 3B-parameter MoE model gains 26% latency reduction with only 0.027% extra parameters, enabling real-time edge deployment.
What the Improved AI System Can Do:
-
Run large MoE models (up to 16B parameters) on memory-constrained edge devices with near-optimal latency and energy efficiency.
-
Dynamically adapt prefetching behavior per token, balancing accuracy and resource use in real time.
-
Guarantee exact model outputs when needed, or gracefully degrade to stall-free operation with negligible quality loss.
-
Operate across diverse hardware (bandwidth, datatype, compute) without manual tuning, making edge AI more accessible and sustainable.
Abstract
Mixture-of-Experts (MoE) models are attractive for edge deployment because they provide high model capacity while activating only a small subset of parameters per token, improving compute efficiency. However, MoE inference at the edge is fundamentally limited by memory. Expert parameters are large and often reside in off-chip memory due to capacity, cost, and power constraints, putting expert loading to the critical path. We present APEX: Adaptive Expert Prefetching, a predictive resource management framework that overlaps expert loading with useful computation. APEX introduces a lightweight prefetch router that predicts candidate experts before the attention block to dynamically fetch additional experts using a learned confidence model. This adaptive strategy achieves over 99% overlap accuracy, significantly outperforming fixed top-k prefetching techniques. APEX supports two execution modes: a correctness-preserving mode that guarantees exact routing semantics, and a stall-free mode that eliminates residual stalls by operating on available experts with negligible impact on application accuracy. Across multiple MoE models, the correctness-preserving mode reduces per-token latency by up to 26% and improves energy-delay product (EDP) by up to 41% over state-of-the-art baselines, while the stall-free mode provides additional efficiency gains with negligible impact on application accuracy. These results establish adaptive, confidence-driven expert prefetching as an effective approach for efficient MoE inference on edge systems.
Sources
- Pre-Attention Expert Prediction and Prefetching for Mixture-of-Experts Large Language Models
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Related papers
- WitCert: Sound Runtime Risk Observability and Gating for KV-Cache Quantization
- Golden Ruler: A Numeric Format Catalog with Bit-Exact Conformance Vectors for FP8, BF16, MXFP4, and Microscaling Formats
- PoisonCap: Efficient Hierarchical Temporal Safety for CHERI
- Provisioning to Runtime Optimization of a 100 MW-Scale AI Cluster
- Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy
- Optimizing Polynomial Multiplication and Fixed-Weight Sampling for HQC on ARM Cortex-M4