Scepsy: Serving Agentic Workflows Using Aggregate LLM Pipelines
cs.DC, cs.AI
Submitted: 2026-04-16
Updated: 2026-09-25
Code: https://github.com/anon/Scepsy
Terminology
Sources
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
- WVA: A Global Optimization Control Plane for llmd
- Artificial Intelligence Index Report 2025
- TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications
- KunServe: Parameter-centric Memory Management for Efficient Memory Overloading Handling in LLM Serving
- DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference
- A Middle Path for On-Premises LLM Deployment: Preserving Privacy Without Sacrificing Model Confidentiality
- Document Ranking with a Pretrained Sequence-to-Sequence Model
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- AIBrix: Towards Scalable, Cost-Effective Large Language Model Inference Infrastructure
- An Empirical Study of Agent Developer Practices in AI Agent Frameworks
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory Ballooning
- JITServe: SLO-aware LLM Serving with Imprecise Request Information
- SeaLLM: Service-Aware and Latency-Optimized Resource Sharing for Large Language Model Inference
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing