GRACE-DS: a Guarded Reward-guided Agent Correction Environment in Data Science
cs.CL, cs.LG
Submitted: 2026-06-14
Updated: 2026-09-11
Code: https://github.com/Alexx221x/GRACE-DS
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce GRACE-DS, a Guarded Reward-guided Agent Correction Environment in Data Science for pre-deployment evaluation of LLM-powered AutoML agents.
Terminology
Abstract
We introduce GRACE-DS, a Guarded Reward-guided Agent Correction Environment in Data Science for pre-deployment evaluation of LLM-powered AutoML agents. GRACE-DS is a set of evaluation metrics in an isolated environment that can be applied to tabular ML tasks specific to a particular organization. It exposes agents to realistic workflow stages, from planning and data inspection through feature engineering, model development, validation, and code repair to final submission, while hidden executable validators measure not only final predictive performance but also leakage avoidance, reproducibility, protocol validity, correction behavior, and reward alignment. The strongest structured regime, flexible iterative interaction (our approach), achieves higher end-to-end normalized hidden-test quality than single-shot generation, unstructured interaction, and restart-based baselines, while also improving protocol-valid completion. Validated across more than 7,000 episodes, these results establish GRACE-DS as a robust platform for assessing the capacity of LLM-based AutoML agents to execute machine learning workflows under production-like conditions and in accordance with organization-specific requirements.
Sources
- AIDE: AI-Driven Exploration in the Space of Code
- DABstep: Data Agent Benchmark for Multi-step Reasoning
- Machine Learning Operations (MLOps): Overview, Definition, and Architecture
- IDA-Bench: Evaluating LLMs on Interactive Guided Data Analysis
- Data Interpreter: An LLM Agent For Data Science
- TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
- DataSciBench: An LLM Agent Benchmark for Data Science
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering