DynBranch: Speculative Subgraph Reuse for Dynamic Agentic LLM Serving
cs.DC, cs.AI, cs.LG
Submitted: 2026-09-25
Updated: 2026-09-25
Code: https://github.com/deepseek-ai/deepseek-harness
Project page: https://deepseek-harness.github.io/deepseek-harness
Terminology
Sources
- Cache-Craft: Managing Chunk-Caches for Efficient Retrieval-Augmented Generation
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
- SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents
- SPORK: Self-Speculative Forking to Accelerate Agentic LLM Inference
- IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents
- AgentScope: A Flexible yet Robust Multi-Agent Platform
- ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
- Dynamic Speculative Agent Planning
- Speculative Interaction Agents: Building Real-Time Agents with Asynchronous I/O and Speculative Tool Calling
- Interactive Speculative Planning: Enhance Agent Efficiency through Co-design of System and User Interface
- Reducing Latency of LLM Search Agent via Speculation-based Algorithm-System Co-Design
- Dense Passage Retrieval for Open-Domain Question Answering
- Measuring Faithfulness in Chain-of-Thought Reasoning
- Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live
- Nightjar: Dynamic Adaptive Speculative Decoding for Large Language Models Serving
- Speculate with Memory: Lossless Acceleration for LLM Agents
- Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
- TurboSpec: Closed-loop Speculation Control System for Optimizing LLM Serving Goodput
- Autellix: An Efficient Serving Engine for LLM Agents as General Programs
Related papers
- iScheduler: Reinforcement Learning-Driven Continual Optimization for Large-Scale Resource Investment Problems
- SAMM: Sharded Automated Market Maker
- InferScale: GPU-Native KV Injection for Personalized LLM Serving
- Vigil: Accountable Liveness against Selective Silence
- Steelhead: Interleaving Partially Synchronous and Asynchronous Commit Rules on a Shared DAG
- Pushing CPU Speech Synthesis to the Wall: Extreme Inference Tuning under Serverless Architecture and Billing