cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents

arXiv:2609.40284 · cs.LG, cs.AI, cs.CL · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents".

Tom: Computer use agents (CUAs), which utilize graphical user interfaces to complete tasks, have recently surpassed human performance on many standard benchmarks,

Jane: First, who's behind it and why it matters.

Paper summary: Jane: So, Tom, if we boil down the main idea of "cua-speedrun," it seems they are building a system that uses a uniform virtual machine setup and an execution pipeline to make sure every comparison is consistent.

Lu: That standardization covers four main dimensions: the agent driven by a large language model, the benchmark defining the task, the environment infrastructure managing VMs and action/observation contracts, and finally, the agent loop that determines interaction.

Meng: So they're trying to decouple things so we can fix two components while only varying one at a time for clean ablation, which makes sense for isolating variables in complex systems.

Lalam: And they aren't just running this on one platform; they use both hosted evaluations through Modal sandboxes and local evaluations using self-hosted vLLM inference servers on fixed L40S GPUs.

Tom: That’s a solid overview, Jane, and what I find particularly interesting is how they address task selection by proposing an energy minimization framework to pick a minimal representative subset of tasks.

Lu: They select a subset K that minimizes the energy distance between the distributions of agent scores on the full benchmark and this smaller set, aiming to minimize end-to-end evaluation time while keeping the signal strong.

Jane: It seems like they're trying to get high statistical power without drowning in evaluation costs, which is a practical concern for any large-scale research project.

Tom: And then they report some really interesting trade-offs in their findings, showing that more reasoning effort can sometimes reduce both task time and model cost because it prevents repeated unsuccessful actions.

Meng: That's a key point for me, Tom; if extra reasoning helps cut down on wasted effort, that speaks directly to how we structure the prompts we give these agents.

Lalam: I think this paper’s focus on turning our attention beyond just performance and success rates is really significant because it forces us to look at reasoning and interaction together when making agents faster.

Conclusion: Tom: So, wrapping up this discussion on "cua-speedrun: Standardized Benchmarking of the Speed of Computer-Use Agents," the authors are essentially saying they've created a reliable way to measure agent speed and efficiency by standardizing the infrastructure and task sets.

Jane: And their conclusion really centers on how this standardization helps us see a clearer picture of the performance-time or performance-cost Pareto frontier when comparing different AI models across various benchmarks.

Lu: The implication is that we can finally maintain a live ranking of the most cost-effective, fastest, and most capable models by using this standardized evaluation methodology.

Meng: For practical application in building real-world agents, it means we can get clearer advice on when to invest more effort into reasoning versus when to optimize for faster environment interaction.

Lalam: I feel that if we take the insights about context-dependent reasoning and batching actions when intermediate screenshots aren't needed, it could fundamentally change how we design agent workflows in our culture.

Tom: That’s what I was thinking; the paper is pushing us to stop just chasing high success rates and start studying reasoning and interaction as two intertwined factors when developing these agents.

Pranjal Aggarwal, Lawrence Keunho Jang, Sean Welleck Daniel Fried, Ruslan Salakhutdinov, Jing Yu Koh

Carnegie Mellon University

cs.LG, cs.AI, cs.CL

Submitted: 2026-09-30

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 91/100

The gist: Computer use agents (CUAs), which utilize graphical user interfaces to complete tasks, have recently surpassed human performance on many standard benchmarks, but their speed and cost remain

Key concepts

CUA-Speedrun
A standardized platform designed to evaluate Computer Use Agents (CUAs). It uses a uniform virtual machine setup and execution pipeline, along with a common agent interface, ensuring consistent testing across different models and benchmarks.
Standardization Dimensions
Four dimensions used to standardize evaluations: the LLM-driven agent, the task benchmark, the environment infrastructure managing VMs and contracts, and the interaction loop. This decoupling allows researchers to fix some components while varying others for clean comparisons.
Energy Minimization Framework
A method used to select a minimal subset of tasks for evaluation. It chooses a subset K that minimizes the energy distance between the full benchmark distribution and K, aiming to preserve statistical signal while reducing total evaluation time and cost.
Pareto Frontier Shift
The finding that no single model family is optimal for all metrics (speed, cost, capability). The optimal trade-off point shifts depending on which specific benchmark is being used, meaning performance depends heavily on the task environment.

Terminology

Summary

Computer use agents (CUAs), which utilize graphical user interfaces to complete tasks, have recently surpassed human performance on many standard benchmarks, but their speed and cost remain significant barriers to widespread deployment. This work introduces cua-speedrun, a standardized platform that addresses the reproducibility crisis in CUA benchmarking by providing uniform infrastructure and task sets for evaluating agent speed and efficiency across different models.

The gist

cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks.

Standardization of Evaluation

The framework standardizes evaluations across four dimensions: (1) the agent driven by a large language model; (2) the benchmark defining the task; (3) the environment infrastructure managing VMs and action/observation contracts; and (4) the agent loop determining interaction. To ensure decoupling, comparisons fix two components while varying one, allowing for clean ablation. Speed is measured by recording time from when instructions are given until termination or limit is reached, separating agent time from environment operations.

Representative Task Selection

To minimize evaluation cost and time while preserving statistical power, the authors propose an energy minimization framework to select a minimal representative subset of tasks. This involves selecting a subset K that minimizes the energy distance between the distributions of agent scores on the full benchmark and the subset, aiming to minimize end-to-end evaluation time while preserving the signal of the results. They determine an optimal size K by ensuring Spearman correlation thresholds are met for (K-1), (K), and (K+1) to prevent selection based on chance.

Infrastructure and Agent Interface

The platform utilizes a serverless cloud-based provider with a uniform CUA evaluation infrastructure, ensuring all evaluations run on the exact same execution pipeline and the same virtual machine setup. A common agent interface is implemented to allow for portable agent harnesses and implementations, enabling a single implementation to operate across different benchmarks. The system supports both hosted evaluations via Modal sandboxes and local evaluations using self-hosted vLLM inference servers on fixed L40S GPUs.

Key Findings on Performance Trade-offs

The analysis reveals that no single model family is optimal for all three [speed, cost, capability], as the Pareto frontier shifts across benchmarks. Key insights include:

  1. More reasoning effort can reduce both task time and model cost for CUAs because it avoids repeated unsuccessful actions.

  2. Faster environment I/O can paradoxically slow down overall task completion time because models are not trained to interact with desktops in different I/O speed configurations.

  3. The optimal reasoning setting is context-dependent: additional reasoning can be valuable on one (harder) benchmark and unnecessary on another (simpler) benchmark.

  4. Batching actions can reduce task time, but this benefit depends on how well a model plans multi-action sequences; batching alone does not make an agent fast.

Practical Suggestions for Building CUAs

The research suggests practical improvements for developing efficient agents:

Choose a frontier configuration for the required performance and budget.

Tune reasoning effort rather than setting it to an extreme. Too little reasoning can lengthen trajectories, while too much can add time and cost without improving performance.

Batch actions when intermediate screenshots are unnecessary.

The paper concludes that the community should focus on turning our attention beyond performance and success rates, emphasizing the need to study reasoning and interaction together when developing faster agents.

Evaluation Metrics

The evaluation reports mean verifier score, mean task time, and mean model cost per task. The framework also tracks detailed interaction metrics: Interaction turns, generated tokens, and exact completion are defined in Appendix D. Token rates are calculated as generated-token count divided by total agent time, which includes the time spent on model requests and agent-side processing.

Benchmark Coverage

The platform evaluates agents across four different CUA benchmarks: OSWorld-Verified, OSWorld 2.0, CUA-World, and MyPCBench, covering long-horizon tasks and personalized computer use. The selection process ensures that the chosen task sets preserve model rankings by maintaining Spearman correlation thresholds of at least 0.95 for (K-1), (K), and (K+1) task counts.

Validation of Infrastructure

The infrastructure includes CUAAutoDebug, an end-to-end test suite to distinguish between model errors and harness errors, ensuring that the evaluation accurately reflects agent capability rather than execution failures. FastCUA is a new I/O system designed to reduce action-to-observation latency from seconds down to milliseconds.

Model Cost and Token Accounting

Model cost is computed by summing the cost of each task’s model requests, excluding environment hosting and setup time.

Improvements for AI systems

Based on the findings of CUA-SPEEDRUN, here are specific, actionable improvements for developing Computer Use Agents (CUAs) and what these enhanced systems can accomplish:


  1. leungh and reasoning effort optimization:

A CUA should not blindly set reasoning effort to maximum. Instead, use the paper's insight that raising reasoning effort can improve benchmark scores while reducing task completion time.

  • Improvement: Implement an adaptive reasoning setting mechanism where the agent starts at a low or medium setting and dynamically increases it only when performance plateaus or when the model is struggling with repeated unsuccessful actions.

  • Improved System Capability: CUAs will complete complex, long-horizon tasks (e.g., multi-step software configuration) significantly faster and with higher success rates by strategically investing computational effort where it yields the highest return on time reduction, rather than just maximizing token generation.

  1. Intelligent Action Batching:

The paper shows that batching actions is beneficial when intermediate screenshots are unnecessary, reducing total task time and token consumption.

  • Improvement: Develop an action-batching heuristic that analyzes the current task state and environment feedback to decide whether to execute multiple related actions in a single request or wait for an observation. This should be prioritized over maximizing the token rate per step.

  • Improved System Capability: CUAs will exhibit significantly reduced wall-time (up to 50% faster on OSWorld) by minimizing inefficient round trips between the agent and the environment, leading to more practical, real-time interaction capabilities for desktop automation.

  1. Harness and Model Pairing Optimization:

The optimal model/harness combination is not universal; it depends on the benchmark and reasoning setting (e.g., GPT-6 Astra performs better with a direct API harness at high effort).

  • Improvement: Create a dynamic Model Configuration Selector that uses the current benchmark profile to recommend the best harness (Direct API vs. Codex) for any given model and reasoning level, based on pre-computed performance trade-offs.

  • Improved System Capability: The agent will achieve peak efficiency by leveraging specialized infrastructure; for example, when tackling a highly visual task (like those on OSWorld), it will automatically switch to the harness proven to handle that specific environment's I/O latency most effectively.

  1. Environment Latency Mitigation (Fast CUA):

The paper demonstrates that optimizing I/O latency (using FastCUA) can sometimes be detrimental if not handled by the model; current models don't always wait optimally for UI updates.

  • Improvement: Integrate a Fast I/O Mode that aggressively minimizes the time between action and observation, while simultaneously using model-specific predictive modeling to estimate necessary UI update delays.

  • Improved System Capability: CUAs will be highly responsive in dynamic desktop environments, reducing perceived lag during tasks like real-time data monitoring or rapid application switching.

  1. Task Selection via Energy Minimization:

Instead of running entire benchmarks, use the proposed energy minimization framework to select only the most representative subset of tasks for evaluation.

  • Improvement: Implement an automated pre-processing step that uses the energy distance metric (Equation 2) to select a minimal set of tasks that preserves the performance ranking and statistical power across a full benchmark.

  • Improved System Capability: The deployment pipeline will drastically reduce evaluation costs and time by focusing resources on tasks that provide the most informative signal, enabling faster iteration cycles for deploying new CUA models.

  1. Robust Error Detection (CUA-AutoDebug):

Since agent failures can be model errors or harness errors, a robust verification layer is critical.

  • Improvement: Integrate the CUA-AutoDebug testing suite directly into the agent's execution loop to distinguish between incorrect model actions and incorrect environmental execution (harness) in real-time.

  • Improved System Capability: The system will achieve much higher reliability by instantly flagging when an action fails due to a flawed instruction (model error) versus a flawed translation of that instruction into a physical mouse click or keystroke (harness error), leading to faster debugging and more stable task completion.

Abstract

Computer use agents (CUAs), which use graphical user interfaces (GUIs) to complete tasks on a computer, have recently surpassed human performance on many standard benchmarks, including difficult long-horizon tasks. Their capabilities are undoubtedly impressive, however, a key barrier to the widespread adoption and deployment of CUAs remains their speed and cost. Progress towards faster yet capable CUAs requires reliable evaluation of their speed, but many CUA benchmarks currently face a reproducibility crisis. Benchmarks are based on complex infrastructure with varying machine and container configurations that confound the evaluation of the execution speed of CUAs. Towards addressing this gap, we propose cua-speedrun, which introduces standardized infrastructure and task sets, with a focus on evaluating the speed and efficiency of CUAs. cua-speedrun uses a uniform virtual machine setup and execution pipeline, along with a common agent interface that enables single-agent implementations to operate seamlessly across different benchmarks. Across four different CUA benchmarks, we evaluate how reasoning effort, agent harnesses, and environment latency affect performance, speed, and cost. We find no single model family is optimal for all three; none of the open-weight models are on the frontier, and also, unintuitively, for some models increasing the reasoning effort can speed up task completion, while faster environment input-output can slow down overall task completion time. We also demonstrate that we can effectively reduce the evaluation task set of most CUA benchmarks without degrading overall statistical power, allowing for more efficient benchmarking and comparison. We believe cua-speedrun will enable structured progress towards fast, efficient CUAs, unlocking new real-world use cases and applications. All code, infrastructure, and analysis are available at https://cuaspeedrun.com.

Sources

Related papers