DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

arXiv:2608.10366 · cs.AI, cs.CL · Submitted 2026-08-11 · Read on arXiv

Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub, Md Tahmid Rahman Laskar, Shafiq Joty, Enamul Hoque Prince

York University · Nanyang Technological University · Salesforce AI Research

cs.AI, cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/vis-nlp/DSAgentBench

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: The paper introduces DSAgentBench, described as "the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments." The authors argue that

Terminology

Summary

The paper introduces DSAgentBench, described as the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments. The authors argue that while real-world data science "involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments, existing benchmarks lack real-computer interaction and do not evaluate whether agents can execute complete end-to-end data-science workflows in realistic computing environments, failing to capture the multi-stage, multi-tool nature of data-science practice."

The paper highlights a critical gap: "Producing syntactically and functionally correct code is fundamentally different from operating autonomously within real computing environments, which requires navigating file systems, coordinating tools, interpreting errors, and refining analyses based on intermediate outputs. The central research question posed is: Can AI agents perform long-horizon reasoning to autonomously execute end-to-end data-science workflows in real computer environments?"

The benchmark is formally defined as "a set of tasks T = (Ci, Ii, Vi) Ni=1, where Ci is a task configuration defining the initial system setup, Ii is a natural language instruction describing the analytical objective, and Vi is a deterministic Python evaluator. The task configuration defines the initialized operating-system state, including available datasets, file system structure, installed libraries, and applications (e.g., IDEs, terminals, browsers), as well as optional initialization or cleanup procedures. Importantly, DSAgentBench evaluates final outcomes rather than prescribing a fixed workflow path: agents may use any available applications as long as outputs satisfy the deterministic evaluator."

Data was collected from diverse real-world sources including "Kaggle competition files (prioritizing top-downloaded and highly rated examples), popular OpenML datasets used in academic evaluation, SQLite databases reflecting production relational storage, GitHub repositories such as Plotly datasets, and web-accessible data retrieved through API. The datasets exhibit varying structural complexity, ranging from flat single-table and multi-table CSVs to normalized relational schemas requiring multi-table joins."

The design was grounded in both prior literature and empirical analysis. The authors manually analyzed 100 high-ranking Kaggle notebooks from popular and historically influential competitions to identify recurring analytical workflows and question types. Two annotators with over five years of data-science experience independently derived an initial task taxonomy, which we then expanded using LLMs to surface underrepresented patterns.

Tasks are organized into six capability categories: Data acquisition, Exploratory data analysis (EDA), Feature engineering, Modeling, Evaluation and deployment, and Visualization and reporting. Four expert annotators collaboratively developed 275 tasks over a three-month period (approximately 400 working hours). Each task includes a natural language instruction, an executable environment configuration, and a deterministic Python evaluation function for automatic verification. Vision-language models were used only to refine wording and identify edge cases, while all task logic, expected outputs, and evaluators were defined and validated by humans.

Each task underwent structured dual review for quality assurance where "one human annotator (the 'creator') developed the task specification and evaluation logic, while a second annotator (the 'verifier') independently assessed instruction clarity, verified executability by running the task with a baseline agent, and validated evaluator correctness. Initial agreement was 86%, with remaining tasks revised through iterative discussion. All 275 tasks ultimately achieved mutual approval."

Key characteristics include:

  • Complexity: Hard: 47.6% Medium: 46.9% Easy: 5.5%

  • Stage Type: Multi-Stage: 56.7% Single-Stage: 43.3%

  • Modality: Tabular: 95.3% Image: 3.6% Text: 1.1%

  • Source: GitHub: 37.1% Kaggle: 29.8% OpenML: 18.9% SQLite: 7.6% Web: 6.5%

  • Tools Used: Python: 100.0% VS Code: 81.1% Jupyter Notebook: 18.9% Chrome: 10.2%

EDA tasks comprise the largest share at 43.3%, which aligns with industry observations that data scientists spend 40-60% of their time on data preparation and exploration. Difficulty levels are assigned based on the number of required steps (easy: 1–2; medium: 3–4; hard: 5+), analytical sophistication, and tool coordination requirements. Tasks average 4–5 analytical steps.

DSAgentBench extends OSWorld to provide a realistic execution environment for end-to-end analytical workflows. The environment runs Ubuntu OS with Python and common data-science libraries pre-installed, and captures screenshots at a resolution of 1920×1080. Visual Studio Code and Jupyter Notebook are available, and Chrome is included for accessing documentation or external data when needed. The environment supports automated data retrieval through the Kaggle API, OpenML, direct URLs, and SQLite databases.

The environment is formally represented as a tuple (O, A, T) with:

  • Observation Space: Either a 1920×1080 Screenshot of the desktop interface or "an integrated Screenshot + Accessibility Tree (A11y) modality that augments the screenshot with structured UI metadata, including element roles, accessible names, bounding boxes, and interaction states, extracted via AT-SPI"

  • Action Space: "mouse-based GUI interactions (e.g., clicking, dragging, scrolling), keyboard input for text entry and shortcut-based control, and meta actions including WAIT to pause execution and allow rendering, DONE to signal task completion, and FAIL to terminate execution with an error"

All agents operate through a standard GUI action space to evaluate end-to-end desktop competence without relying on privileged task-specific APIs.

The paper evaluates a diverse set of vision-language agents spanning closed-source, hybrid, and open-source architectures:

  • Closed-source: GPT-4o, GPT-5-mini, GPT-5, O4-mini, Claude Sonnet 4, 4.5 and 4.6, Gemini 2.5 Pro, and OpenAI's Computer Agent

  • Hybrid: Jedi-3B and Jedi-7B paired with GPT-4o

  • Open-source: UI-TARS 2B and 7B, GUI-OWL-7B, and OpenCUA-72B

Each task is paired with a deterministic Python evaluator executed after the agent completes its interaction sequence. The evaluator collects relevant artifacts from the virtual machine, including scripts, generated files, visualizations, or trained models, and scores them against task-specific evaluation criteria. For example, "in a correlation analysis task, the evaluator verifies that the required output file exists, loads the computed Pearson correlation coefficient, and checks whether it matches the expected value within a tolerance of ϵ = 0.01."

For visualization tasks, the evaluator checks that a visualization is generated and verifies required metadata such as labeled axes, titles, legends, and correct data mappings. GPT-4o is used as a visual judge, but to avoid evaluation circularity, we use Gemini-2.5-Pro as the visual judge for all GPT-4o generated outputs, while GPT-4o serves as the judge for outputs from other models. Only ∼10% of tasks invoke an LLM-based judge, and only after deterministic validation gates.

The primary metric is task success rate, defined as the percentage of tasks achieving an overall score greater than or equal to 0.95.

Three participants, two applied scientists and one master's graduate, achieved an overall success rate of 85.09% under the same task environment and deterministic evaluation protocol.

The results reveal a clear performance gap between current agents and the demands of real-world end-to-end data science workflows:

  • Claude-4.6-Sonnet achieves the strongest agent performance at 56.70% overall accuracy under the Screenshot + A11y Tree setting

  • GPT-5 follows at 29.81%

  • Other closed-source models (GPT-4o, Gemini-2.5-Pro, GPT-5-mini) achieve around 20%

  • Open-source agents achieve at most 1% accuracy under the Screenshot-only setting, and none support A11y Tree input

The paper notes: Despite generating code and issue interface actions, these models consistently fail to ground instructions in the UI state and complete end-to-end workflows.

Data acquisition, model validation, and evaluation tasks remain most challenging, reflecting the difficulty of long-horizon reasoning and tool coordination. Adding accessibility information generally improves performance, indicating that structured UI metadata aids grounding and interaction, though gains vary widely across models.

Single-stage tasks achieve substantially higher success rates than multi-stage workflows, highlighting the difficulty of maintaining state, recovering from errors, and coordinating tools across longer execution chains. Performance degrades monotonically with task difficulty.

Increasing the budget from 15 to 50 steps yields only marginal gains in task success rate (24.54% → 25.81%) and average score (0.55 → 0.57). The paper concludes: performance is not primarily limited by the action budget; instead, failures reflect a combination of grounding, planning, reasoning, tool orchestration, and analytical execution errors.

Prompting GPT-4o to prefer terminal or command-line workflows whenever they could solve the current subtask more directly and reliably changed performance only marginally, from 19.34% to 20.73%, suggesting that terminal-first execution alone does not substantially reduce the benchmark difficulty.

The analysis of 604 manually inspected runs from closed-source models using Screenshot + A11y Tree observations, along with 150 sampled runs from open-source models reveals:

  • Open-source agents fail almost entirely due to grounding errors (97–98%), revealing poor alignment with desktop state

  • GPT-4o still fails mainly from grounding

  • Better-grounded models (Gemini-2.5-Pro, GPT-5, Claude-4.6-Sonnet) exhibit terminal, code, and reasoning failures

CUA and GUI-OWL-7B exhibit predominantly late-stage failures, with over 93% occurring after prolonged interaction, reflecting extended but ineffective exploration due to poor grounding. In contrast, GPT-4o and Gemini-2.5-Pro show a higher proportion of early and mid-trajectory failures, indicating weaker robustness in initial grounding and planning. Open-source models fail especially early, often unable to correctly open or control the terminal.

Gemini-2.5-Pro completes tasks most efficiently (6.76 steps on average), followed by GPT-5-Mini (7.33), GPT-4o (10.03), and CUA (15.00). This reveals a clear trade-off between exploration depth and execution efficiency: models that act conservatively and terminate earlier achieve faster successes but also incur higher early-failure rates.

The paper identifies recurring failure patterns that are consistent across model families:

  1. Difficulty in handling environmental notifications and system-level UI elements - models frequently fail to correctly interpret or dismiss generic pop-up notifications originating from the code editor environment

  2. Systematic weaknesses in maintaining proper code formatting and indentation - models intend to insert a line break but instead explicitly generate the string token within the code, resulting in syntactically invalid or semantically incorrect code

The paper summarizes four main contributions:

  1. DSAgentBench, the first benchmark for evaluating autonomous data-science workflows inside real operating systems, covering the full data-science lifecycle

  2. An extension of OSWorld that enables interaction with core data-science tools and external data sources (e.g., Kaggle, OpenML, SQLite)

  3. "Deterministic, execution-based evaluation that assesses analytical correctness, visualization outputs, and model performance beyond code-only execution, along with an extensive evaluation of 15 open- and closed-source agents"

  4. In-depth analysis and ablations that identify key limitations in grounding, planning, and analytical reasoning and outline directions for future agent development

The paper acknowledges several limitations:

  1. Open-source agents in our evaluation stack do not currently support A11y Tree observations, so we evaluate them under the screenshot-only setting

  2. Our detailed error analysis is based on 604 manually inspected trajectories from closed-source models and 150 from open-source models... it still represents a subset of the full benchmark

  3. Our evaluation of visualization-heavy tasks focuses on final artifact quality... overall visual clarity and semantic alignment are assessed through the LLM judge

The paper concludes: "We present DSAgentBench, the first benchmark for assessing whether agents can automate end-to-end data-science workflows inside real operating systems. Unlike prior benchmarks that focus on isolated code generation or generic GUI interaction, our benchmark requires agents to plan and execute long-horizon, multi-tool workflows spanning data acquisition, analysis, and modeling. Our experiments reveal substantial limitations in current agents, with even the strongest systems achieving low success rates on these tasks." The benchmark is released at https://github.com/vis-nlp/DSAgentBench.

Improvements for AI systems

Based on the paper, here are specific improvements for AI systems:

  • Improvement: Train agents to fuse screenshot pixels with structured accessibility tree data (element roles, names, bounding boxes) rather than relying on vision alone. The paper shows A11y integration boosts performance (e.g., Claude-4.6-Sonnet reaches 56.7% vs. 20% for vision-only models).

  • Capability: Agents can accurately locate and interact with UI elements (buttons, menus, file dialogs) in desktop environments, reducing grounding errors that cause 97–98% of open-source agent failures.

  • Improvement: Implement a two-level planner: (a) a high-level task decomposer that breaks multi-stage workflows (data acquisition → EDA → modeling → validation) into sub-goals, and (b) a low-level executor that tracks intermediate artifacts (files, variables, plots) and can backtrack on failure. The paper shows single-stage tasks succeed at much higher rates than multi-stage ones.

  • Capability: Agents maintain context across 5+ analytical steps, recover from errors (e.g., failed file writes, incorrect joins), and re-plan without restarting, addressing the 56.7% multi-stage task failure.

  • Improvement: Add a dedicated code-execution module that: (a) validates syntax before insertion, (b) handles indentation and line breaks natively (not as string tokens), and (c) interprets terminal output (errors, warnings) to self-correct. The paper identifies formatting errors and terminal misreading as recurring failures.

  • Capability: Agents write syntactically correct Python, execute shell commands reliably (e.g., pip install, git clone), and debug runtime errors autonomously, reducing the 20–30% of failures from code/terminal issues in better-grounded models.

  • Improvement: Train agents to validate intermediate outputs (e.g., check file existence, data shape, column names) after each major step, rather than exploring extensively before acting. The paper shows successful runs average 6.76–15 steps, but failures often occur late due to ineffective exploration.

  • Capability: Agents detect errors early (e.g., wrong CSV loaded, missing join key) and correct course, improving success rates on hard tasks (47.6% of benchmark) that require 5+ steps.

  • Improvement: Build a persistent workspace memory that tracks: (a) which tools (VS Code, Jupyter, Chrome, terminal) are open, (b) current file paths and data schemas, and (c) previously executed commands. This addresses the paper's finding that models struggle to coordinate notebooks, terminals, and browsers.

  • Capability: Agents seamlessly switch between writing code in VS Code, querying SQLite, fetching data via Chrome, and running validation scripts—without losing state or re-discovering the environment.

  • Improvement: Add a dedicated sub-module to recognize and dismiss non-task-critical UI elements (pop-ups, update prompts, editor notifications) using both visual patterns and accessibility metadata. The paper notes models frequently fail on these.

  • Capability: Agents ignore irrelevant distractions and focus on analytical tasks, reducing early failures and wasted steps.

  • Improvement: Implement a confidence estimator that predicts task completion probability after each action. If confidence is low after 15–20 steps, trigger a re-planning phase (e.g., re-read instructions, inspect file system) rather than continuing blindly. The paper shows increasing step budget from 15 to 50 yields only 1.27% improvement.

  • Capability: Agents allocate effort efficiently—terminating early on simple tasks (1–2 steps) and investing more on hard tasks only when progress is detected, avoiding the long ineffective exploration failure mode.

  • Improvement: Add a post-generation validator that checks: (a) axis labels, titles, legends exist, (b) data mappings are correct (e.g., x-axis is date, y-axis is sales), and (c) plot type matches the task (e.g., scatter for correlation). Use deterministic checks first, LLM judge only as fallback.

  • Capability: Agents produce publication-ready visualizations that pass automated quality gates, reducing the need for subjective LLM evaluation and improving reproducibility.


The enhanced system can autonomously:

  • Acquire data from Kaggle, OpenML, SQLite, or web APIs by navigating real file systems and browsers.

  • Perform full EDA (missing value analysis, correlation matrices, distribution plots) in VS Code or Jupyter, with correct code formatting and error recovery.

  • Engineer features (e.g., date parsing, one-hot encoding, aggregations) across multi-table relational schemas.

  • Train and validate models (e.g., scikit-learn pipelines) with proper train/test splits and metric computation.

  • Generate and verify visualizations with correct labels, legends, and data mappings.

  • Coordinate multiple tools (terminal for package installs, VS Code for scripting, Chrome for documentation) without losing context.

  • Self-correct by reading terminal errors, checking intermediate files, and re-planning when stuck.

This would raise success rates from the current best of 56.7% toward the human baseline of 85.09%, particularly on hard, multi-stage tasks involving data acquisition, model validation, and evaluation.

Abstract

Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, terminals, browsers, and databases within real operating environments. Yet existing benchmarks lack real-computer interaction and do not evaluate whether agents can execute complete end-to-end data-science workflows in realistic computing environments, failing to capture the multi-stage, multi-tool nature of data-science practice. We introduce DSAgentBench, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments. DSAgentBench contains 275 diverse tasks covering the entire data-science life-cycle, reflecting the complexity and tool coordination required in practice. Each task requires grounding decisions in intermediate outputs and coordinated tool use, and includes a deterministic evaluator that verifies analytical correctness, visual outputs, and model performance rather than code-only execution. Our extensive experiments with 15 closed- and open-source models show that even the strongest agent, Claude-4.6-Sonnet, achieves only 56.70% task success, while all open-source agents remain below 1%, frequently failing at tool orchestration, OS grounding, and multi-step reasoning. These results reveal a substantial capability gap between current agentic systems and real data-science workflows, positioning DSAgentBench as a foundation for developing grounded, verifiable, autonomous data-science agents. We release DSAgentBench at https://github.com/vis-nlp/DSAgentBench.

Sources

Related papers