TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Terminology
Sources
- Self-Resource Allocation in Multi-Agent LLM Systems
- Kairos: Low-latency Multi-Agent Serving with Shared LLMs and Excessive Loads in the Public Cloud
- MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- DroidSpeak: KV Cache Sharing for Cross-LLM Communication and Multi-LLM Serving
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- Astraea: A State-Aware Scheduling Engine for LLM-Powered Agents
- Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
- Batch Query Processing and Optimization for Agentic Workflows
- FlowMesh: A Service Fabric for Composable LLM Workflows
- Preble: Efficient Distributed Prompt Scheduling for LLM Serving
- Teola: Towards End-to-End Optimization of LLM-based Applications
- Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference Serving
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering