Crossflow: Prefill-Decode Elasticity for Agentic LLM Serving
cs.DC, cs.AI, cs.LG
Submitted: 2026-09-22
Updated: 2026-09-22
Terminology
Sources
- DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
- Punica: Multi-Tenant LoRA Serving
- KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
- LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
- GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention
- FLYING SERVING: On-the-Fly Parallelism Switching for Large Language Model Serving
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- M'elange: Cost Efficient Large Language Model Serving by Exploiting GPU Heterogeneity
- DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
- MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
- AdaSpec: Adaptive Speculative Decoding for Fast, SLO-Aware Large Language Model Serving
- Hydragen: High-Throughput LLM Inference with Shared Prefixes
- ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System
- Fast Inference from Transformers via Speculative Decoding
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Not All Prefills Are Equal: PPD Disaggregation for Multi-turn LLM Serving
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing