HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Xin Yuan, Ning Li, Wenchao Xu, Song Guo, Haijun Zhang
Harbin Institute of Technology · University of Science and Technology Beijing · Hong Kong University of Science and Technology
cs.NI, cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 15 pages, 9 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: HetRoute: Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference Summary This paper proposes HetRoute, a heterogeneous and cost-aware collaborative routing
Terminology
Summary
HetRoute: Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference
Summary
This paper proposes HetRoute, a heterogeneous and cost-aware collaborative routing framework for distributed edge Mixture-of-Experts (MoE) inference. The framework addresses the challenge of deploying large language models (LLMs) with MoE architecture across geo-distributed, heterogeneous edge servers connected by ordinary Internet links.
Problem Statement and Motivation
The paper identifies that when the Top-k activated experts of a token are spread across multiple servers, the optimal routing depends jointly on four factors: cross-server link bandwidth, heterogeneous GPU computing capability, GPU-CPU expert loading delay, instantaneous queueing backlog, and replica-level quantization quality loss. Existing distributed inference and MoE serving methods address these factors separately and do not provide a unified framework for online multi-server collaborative routing.
The authors highlight several new challenges that arise in distributed collaborative MoE inference over heterogeneous edge servers: (1) Top-k activated experts of one token may be distributed over multiple edge servers, making token routing a multi-server collaborative routing problem rather than a single-server selection problem; (2) experts are deployed cooperatively between GPU and CPU inside each server, where hot experts are kept in GPU memory and cold experts in CPU memory, meaning a locally available cold expert may incur significant GPU-CPU loading delay and may be slower than a remote hot expert; (3) inter-server links in edge networks are ordinary and heterogeneous Internet links rather than stable high-speed data center connections; (4) different replicas may be stored with different quantization precisions, introducing different quality losses.
Proposed Framework
HetRoute introduces a unified per-assignment cost model that explicitly captures four cost components: cross-server transmission, GPU-CPU offloading, GPU computation with queueing, and quantization-induced quality penalty. The framework operates in two stages:
Offline Stage (Algorithm 1): Determines expert server placement, GPU-CPU residency, and replica precision through a routing-cost-coupled deployment algorithm. The offline stage uses calibration traffic and the online unified cost to estimate replica-selection probabilities, GPU-residency benefits, and redundancy benefits. It consists of three stages: (1) routing-probability-weighted GPU residency, which promotes replicas to GPU memory based on expected offload latency savings weighted by online selection probability; (2) unified-cost-driven redundant replication, which adds replicas when the expected reduction in online unified cost exceeds a memory price; (3) global residency re-optimization with a diminishing-returns guard.
Online Stage (Algorithm 2): Routes the Top-k activated expert set as a whole by minimizing the bottleneck layer cost via exact enumeration or beam search. The algorithm constructs a candidate collaboration domain for each Top-k activated expert set, prunes candidates according to quality and stability guards, and selects a complete collaboration set by minimizing the layer-level bottleneck cost. The online router uses an exact-first substitute policy, where substitute execution is only activated when no exact candidate remains admissible. An emergency fallback mechanism routes to full-precision exact replicas guaranteed by constraint (c.4), which is always quality-safe.
Key Technical Contributions
-
Unified cost model: The per-assignment cost function captures transmission, offloading, computation with queueing, and quality penalty, enabling the framework to recognize that a remote GPU-resident hot expert may be faster than a local CPU-resident cold expert.
-
Set-level collaborative routing: Instead of greedily selecting the best server for each expert independently, HetRoute evaluates the bottleneck delay of the complete collaboration set, including fan-out transmission, parallel expert execution, fan-in aggregation, GPU-CPU loading delay, and quality cost.
-
Offline-online co-optimization: The offline deployment is coupled with the online routing objective, where GPU residency benefit and redundant-replication benefit are evaluated according to the expected online unified cost under calibration traffic.
-
Theoretical properties: The paper establishes fallback feasibility (Property 1), a bound on the number of participating servers (Property 2), per-layer optimality for small candidate domains (Property 3), local optimality of the exchange refinement (Property 4), and online computational complexity of O(LkRB) for beam-search and O(LR k) for exact variants (Property 5).
Experimental Results
The framework is evaluated on a trace-driven heterogeneous 10-server edge testbed using three MoE models: Switch-Base-8E, Qwen-MoE-A 2.7B, and Mixtral-8x7B. The servers have GPU computing capacities spanning 18 to 120 TFLOPS, GPU memory capacities of 12 to 48 GB, CPU memory capacities of 96 to 512 GB, GPU-CPU transfer bandwidths of 12 to 64 GB/s, and inter-server bandwidths of 0.8 to 10 Gbps.
Key results on Mixtral-8x7B include:
-
Average latency reduction of up to 59.0% compared with EdgeShard, 49.5% lower than Petals, 38.6% lower than MoE-Infinity, and 28.1% lower than Prism (achieving 156 ms average latency)
-
P99 latency reduction of up to 58.0% (286 ms, 33.2% lower than Prism)
-
Cross-server traffic reduction of up to 72.1% (1.24 GB per one thousand tokens, 24.8% below Prism)
-
2.13x throughput improvement compared with EdgeShard and 1.35x compared with Prism (1490 tokens per second)
-
Remote execution ratio of 22% versus 68% for EdgeShard and 61% for Petals
-
CPU-offload ratio reduced to 14% compared with 42% for MoE-Infinity and 31% for Prism
The paper also demonstrates that quality degradation remains within the configured budget (2% default), with perplexity on WikiText-103 increasing from 5.32 to 5.39, SQuAD F1 dropping from 86.7 to 86.1, and GSM8K accuracy dropping from 56.8% to 56.0% compared with the centralized full-precision model. The quality budget can be traded for latency in a smooth and monotone manner, with the knee of the curve appearing around the 2% budget.
Ablation studies show that removing set-level routing causes the largest latency increase (from 156 ms to 203 ms), removing GPU-CPU residency optimization increases latency to 218 ms and P99 latency to 421 ms, removing quality-aware precision selection increases latency to 181 ms, and removing redundant replication increases latency to 194 ms. Under heavy load (80 requests per second), HetRoute satisfies 73% of requests below the 300 ms SLA target, compared with 51% for Prism, 43% for MoE-Infinity, and 21% for EdgeShard.
Improvements for AI systems
Improvements to AI Systems:
-
Unified Cost-Aware Routing for Distributed Inference: Implement HetRoute’s per-assignment cost model (transmission, GPU-CPU offloading, queueing, quantization penalty) into any distributed LLM serving system. This enables the AI to dynamically choose between local CPU-resident experts and remote GPU-resident experts based on real-time network bandwidth, GPU load, and memory transfer speeds, rather than assuming local execution is always optimal. The improved system can reduce average latency by up to 59% in heterogeneous edge environments.
-
Set-Level Collaborative Routing for Multi-Expert Tokens: Replace greedy per-expert server selection with HetRoute’s bottleneck-layer optimization (exact enumeration or beam search). This allows the AI to jointly decide the entire Top-k expert placement for a token, minimizing the slowest stage (fan-out, parallel compute, fan-in) across all servers. The improved system can cut P99 latency by 58% and reduce cross-server traffic by 72%, enabling real-time MoE inference over ordinary Internet links.
-
Offline–Online Co-Optimized Deployment: Use HetRoute’s offline algorithm to pre-place experts on GPU/CPU and choose quantization precision based on predicted online routing costs (from calibration traffic). This lets the AI system proactively store hot experts in GPU memory and replicate them with appropriate precision, reducing CPU-offload ratio from 42% to 14% and improving throughput by 2.13x under load, while keeping quality degradation within a user-defined budget (e.g., 2% perplexity increase).
-
Quality-Aware Precision Selection with Fallback Guarantees: Integrate HetRoute’s replica-level quantization selection and emergency fallback to full-precision exact replicas. The AI can trade off latency for quality smoothly (e.g., 2% budget yields 156 ms latency; larger budgets yield faster responses) and automatically route to a safe, full-precision replica when no admissible low-quality candidate exists. This ensures the system never violates a hard quality constraint, even under network instability or server failures.
-
Queueing-Aware Load Balancing: Incorporate HetRoute’s instantaneous queueing backlog into routing decisions. The improved system can avoid sending tokens to overloaded servers, even if they have high GPU capacity, by factoring in current execution queues. This enables the AI to satisfy 73% of requests under a 300 ms SLA at 80 requests per second (vs. 51% for state-of-the-art baselines), making it suitable for latency-sensitive edge applications like interactive chatbots or real-time code assistants.
-
Heterogeneous Hardware Adaptation: Leverage HetRoute’s cost model to automatically adapt to servers with wildly different GPU TFLOPS (18–120), memory (12–48 GB GPU, 96–512 GB CPU), and link speeds (0.8–10 Gbps). The improved AI system can run on a mixed fleet of edge devices without manual tuning, balancing load across weak and strong nodes to maximize aggregate throughput (1490 tokens/sec on Mixtral-8x7B) while minimizing remote execution ratio (22% vs. 68% for naive sharding).
Abstract
Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous edge servers remains challenging. When the Top-k activated experts of a token are spread across multiple servers, the optimal routing depends jointly on cross-server link bandwidth, heterogeneous GPU computing capability, GPU-CPU expert loading delay, instantaneous queueing backlog, and replica-level quantization quality loss. Existing distributed inference and MoE serving methods address these factors separately and do not provide a unified framework for online multi-server collaborative routing. In this paper, we propose HetRoute, a heterogeneous-cost-aware collaborative routing framework for distributed edge MoE inference. HetRoute introduces a unified per-assignment cost model that explicitly captures four cost components: cross-server transmission, GPU-CPU offloading, GPU computation with queueing, and quantization-induced quality penalty. Guided by this model, the offline stage determines expert server placement, GPU-CPU residency, and replica precision through a routing-cost-coupled deployment algorithm, while the online stage routes the Top-k activated expert set as a whole by minimizing the bottleneck layer cost via exact enumeration or beam search. Theoretical analysis establishes fallback feasibility, a bound on the number of participating servers, per-layer optimality for small candidate domains, and online computational complexity. Trace-driven evaluation on three MoE models over a heterogeneous 10-server edge testbed shows that HetRoute reduces average inference latency by up to 59.0% and P99 latency by up to 58.0%, cuts cross-server traffic by up to 72.1%, and achieves 2.13x throughput improvement compared with representative baselines, while keeping quality degradation within the configured budget.
Sources
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- GPT-4 Technical Report
- Mixtral of Experts
- MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
- Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
- OrderMoE: An expert similarity driven distributed edge MoE inference
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic