EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
cs.DC, cs.LG, cs.PF
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/fixie-ai/ultravox
Terminology
Sources
- Yi: Open Foundation Models by 01.AI
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- Qwen2.5-VL Technical Report
- HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
- Gemma 2: Improving Open Language Models at a Practical Size
- Niyama : Breaking the Silos of LLM Inference Serving
- semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage
- LLaVA-OneVision: Easy Visual Task Transfer
- Nexus:Proactive Intra-GPU Disaggregation of Prefill and Decode in LLM Serving
- Fast Distributed Inference Serving for Large Language Models
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing