EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
cs.DC, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/KunmingSHAO/efficientagent_release
Terminology
Sources
- SkyRL-Agent: Efficient RL Training for Multi-turn LLM Agent
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
- TOPAS: Workflow-Aware Prefix-State Scheduling for Multi-Agent LLM Serving
- Scaling Long-Horizon LLM Agent via Context-Folding
- SWE-Pruner: Self-Adaptive Context Pruning for Coding Agents
- Idleness is Relative: Exploiting Tool-Call Idle Windows for Offloading in Agentic Systems with MORI
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing