TopoCompress: Topology Aware Token Compression Algorithm for Distributed Edge MoE Inference
cs.NI, cs.CL
Submitted: 2026-08-01
Updated: 2026-09-23
Comments: 15 pages, 9 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token.
Terminology
Abstract
Mixture-of-experts (MoE) models improve capacity with moderate overhead by sparsely activating experts per token. However, deploying MoE across resource-constrained edge servers incurs substantial cross-server communication as experts are distributed across heterogeneous servers. Existing placement methods optimize for raw token traffic, while conventional compression considers semantics but ignores topology-dependent routing costs. Consequently, independent optimization leads to inefficient communication and resource utilization. This paper proposes TopoCompress, a deployment- and topology-aware token compression framework for communication-efficient distributed edge MoE inference. It jointly optimizes token compression, expert deployment/replication, GPU-CPU residency, and collaborative routing to balance cross-server transmission, quality, and resource use. To address the coupling between token-level compression and epoch-level deployment, TopoCompress employs a two-timescale alternating optimization. In the online fast loop, it identifies and compresses low-importance, high-routing-cost tokens and jointly routes surviving expert activations. In the offline slow loop, it updates expert placement, replication, and GPU-CPU residency according to post-compression traffic accumulated during online inference. We establish the feasibility, optimality, convergence, and computational complexity. Simulations demonstrate that TopoCompress effectively reduces cross-server traffic and deployment resource consumption while maintaining controllable inference quality, enabling efficient distributed MoE inference over bandwidth- and resource-constrained edge infrastructures.
Sources
- A Survey of Large Language Models
- A Survey on Multimodal Large Language Models
- Mixtral of Experts
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- A Survey on Efficient Inference for Large Language Models
- Model-Distributed Inference for Large Language Models at the Edge
- MoE-Infinity: Efficient MoE Inference on Personal Machines with Sparsity-Aware Expert Cache
- Accelerating Edge Inference for Distributed MoE Models with Latency-Optimized Expert Placement
- EdgeMoE: Empowering Sparse Large Language Models on Mobile Devices
- SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization
- PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
- MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs
- OrderMoE: An expert similarity driven distributed edge MoE inference
Related papers
- HiFiNet: Hierarchical Fault Identification in Wireless Sensor Networks via Edge-Based Classification and Graph Aggregation
- Embodied AI in 6G Networks: From Intelligent Connectivity to Physical Intelligence
- Lightweight GenAI for Network Traffic Generation: Fidelity, Augmentation, and Classification
- EdgePoW: Adaptive Ingress-Aware Defense with Non-Interactive PoW Against Volumetric SYN Floods
- SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks
- What is Normal? A Big Data Observational Science Model of Anonymized Internet Traffic