The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction
cs.AI
Submitted: 2026-09-16
Updated: 2026-09-17
Code: https://github.com/Edge0-AI/Edge0
License: http://creativecommons.org/licenses/by/4.0/
The gist: Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held.
Terminology
Abstract
Mixture-of-experts (MoE) inference on consumer hardware is bounded by weight memory: a 35B-class model is 19.5GB at 4-bit, and sparsity shrinks the compute per token, not the bytes that must be held. Naive offloading to SSD does not help on its own, because layer N+1's experts must be chosen before layer N's output exists, so the reads cannot start early enough to hide behind compute. We present Edge0, a streaming MoE inference engine that closes the gap with a prerouter: a per-layer head predicts the next layer's routing one token ahead, and the prediction is consumed as the routing itself, so the staged expert set equals the routed set and nothing is dropped. An unmerged recovery LoRA, trained on the student path, pays back the quality lost to int4 quantization and routing replacement. On a single 24GB machine, Edge0 serves a 35B MoE at 20tok/s inside 3GiB of peak active memory, within a few points of its fp16 teacher on average across five public benchmarks. An 8B tier runs on the same framework, and the framework, checkpoints, and adapters are open source.
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- Tiny-Engram: Trigger-Indexed Concept Tables for Generative Vision
- Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
- DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
- DeepSeek-V3 Technical Report
- QLoRA: Efficient Finetuning of Quantized LLMs
- Fast Inference of Mixture-of-Experts Language Models with Offloading
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- MiniLLM: On-Policy Distillation of Large Language Models
- Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
- Mixtral of Experts
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Post-Trained MoE Can Skip Half Experts via Self-Distillation
- DeepSpeed-MoE: Advancing Mixture-of-Experts Inference and Training to Power Next-Generation AI Scale
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
- PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU
- Every FLOP Counts: Scaling a 300B Mixture-of-Experts LING LLM without Premium GPUs
- MultiLoRA: Democratizing LoRA for Better Multi-Task Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection