SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading
cs.OS, cs.AI
Submitted: 2026-09-11
Updated: 2026-09-11
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Related papers
- MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices
- Planarian: Managing Agent State with Statepoints
- VUDA: Enabling Controlled Spatial Sharing of Graphics and Compute on NVIDIA GPUs
- ActKV: Efficient LLM Agents through Action-Guided KV Cache Management
- GroupKV: Hierarchical KV Cache Management for Long-Context Diffusion LLM Inference
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live