Handoff Debt: The Rediscovery Cost When Coding Agents Take Over Interrupted Tasks
cs.AI
Submitted: 2026-06-01
Updated: 2026-08-30
Code: https://github.com/crewAIInc/crewAI
License: http://creativecommons.org/licenses/by/4.0/
The gist: Coding-agent benchmarks evaluate whether a single uninterrupted agent can resolve a repository issue.
Terminology
Abstract
Coding-agent benchmarks evaluate whether a single uninterrupted agent can resolve a repository issue. Real software work is messier: tasks are interrupted, reassigned, reviewed, and resumed from partial states left by another agent or engineer. We study this missing dimension through handoff debt: the rediscovery cost imposed when a predecessor's work is opaque or incomplete. Our takeover protocol interrupts a coding agent at deterministic handoff points, freezes the repository, and evaluates successor agents under four handoff views: repository state only, raw trace, summary notes, and structured notes. Across 75 source tasks, the protocol generates 181 handoff-point tasks and 724 takeover runs per successor model. Across three successor models, context-bearing handoffs reduce median agent events by 20-59% and cumulative prompt tokens by 42-63% relative to repository-only takeover. Solved-rate effects are smaller and model-dependent, but efficiency gains are consistent. These findings suggest that coding-agent evaluation should report not only whether a task is solved, but also how costly that work is for another agent to resume.
Sources
- Longformer: The Long-Document Transformer
- Agyn: A Multi-Agent System for Team-Based Autonomous Software Engineering
- ContextBench: A Benchmark for Context Retrieval in Coding Agents
- From Guessing to Seeing: Enhancing LLM-Based Program Repair via Trace-Guided Multi-strategy Debate
- AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation
- REFINE: Enhancing Program Repair Agents through Context-Aware Patch Refinement
- Agentless: Demystifying LLM-based Software Engineering Agents
- MemGPT: Towards LLMs as Operating Systems
- $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
- HAFixAgent: History-Aware Program Repair Agent
- Voyager: An Open-Ended Embodied Agent with Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection