CUADebug: Diagnosing and Repairing Computer-Use Agent Failures

arXiv:2608.02643 · cs.SE, cs.AI · Submitted 2026-07-31 · Read on arXiv

cs.SE, cs.AI

Submitted: 2026-07-31

Updated: 2026-09-06

Comments: 23 pages, 10 figures, 6 tables

License: http://creativecommons.org/licenses/by/4.0/

The gist: Computer-use agents (CUAs) interact with graphical interfaces through screenshots and low-level mouse and keyboard actions, yet the causal error may precede the terminal failure.

Terminology

Abstract

Computer-use agents (CUAs) interact with graphical interfaces through screenshots and low-level mouse and keyboard actions, yet the causal error may precede the terminal failure. We present CUADebug, a framework for localizing root causes in CUA trajectories and guiding re-execution. CUADebug includes a five-category, 30-subtype taxonomy; CUAErrorBench, a benchmark of 204 failed OSWorld trajectories with human root-cause annotations; and CUADebugger, a ReAct-style agent for root-cause analysis (RCA). CUADebugger iteratively selects trajectory steps, inspects paired before/after screenshots and action traces, and submits a structured diagnosis containing the causal step, taxonomy label, grounded evidence, and correction. CUADebugger performs RCA without per-trajectory human intervention; human annotations are used to evaluate RCA predictions and, in controlled re-rollout comparisons, to fix restart points. Task reasoning and control is the largest annotated failure category (110/204). CUADebugger improves L2 and Tag+Step Exact across three debugger backbones on the Claude-agent split; with Gemini 2.5 Pro, Tag+Step Exact rises from 11.1% to 19.4%. Single re-execution improves failure recovery from 13.89% to 29.86% (overall 61.77% to 68.14%); controlled continual re-execution improves it from 12.50% to 25.69% (overall 61.22% to 66.48%). Project page: cuadebug.github.io.

Sources

Related papers