Do Not Restart: Residual Completion for Stateful Agent Handoffs
cs.AI
Submitted: 2026-09-12
Updated: 2026-09-26
Comments: 11 pages, 2 figures, 4 tables
Code: https://github.com/microsoft/STATE-Bench
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- From Confident Closing to Silent Failure: Characterizing False Success in LLM Agents
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Is Escalation Worth It? On the Depth of LLM Cascades
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- Cordon: Semantic Transactions for Tool-Using LLM Agents
- Defeating Prompt Injections by Design
- The Granularity Mismatch in Agent Security: Argument-Level Provenance Solves Enforcement and Isolates the LLM Reasoning Bottleneck
- Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation
- Enforcing Temporal Constraints for LLM Agents
- Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks
- VIGIL: Defending LLM Agents Against Tool Stream Injection via Verify-Before-Commit
- Agentic Routing: The Harness-Native Data Flywheel
- ToolGate: Contract-Grounded and Verified Tool Execution for LLMs
- ToolSandbox: A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
- RouteLLM: Learning to Route LLMs with Preference Data
- GoEX: Perspectives and Designs Towards a Runtime for Autonomous LLM Applications
- Agent-Diff: Benchmarking LLM Agents on Enterprise API Tasks via Code Execution with State-Diff-Based Evaluation
- SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks
- Agent Workflow Memory
- Step-level Optimization for Efficient Computer-use Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection