Skill-based Agentic Evaluation for Real-time Data Science Tasks
cs.AI, cs.LG, cs.MA
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring.
Terminology
Abstract
We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: "what were last week's audience sizes"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.
Sources
- DS-STAR: Data Science Agent for Solving Diverse Tasks across Heterogeneous Formats and Open-Ended Queries
- A Survey on Large Language Model-based Agents for Statistics and Data Science
- Large Language Model-based Data Science Agent: A Survey
- Measuring Data Science Automation: A Survey of Evaluation Tools for AI Assistants and Agents
- LiveBench: A Challenging, Contamination-Limited LLM Benchmark
- Natural-Language Agent Harnesses
- SoK: Agentic Skills -- Beyond Tool Use in LLM Agents
- DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning
- Data Interpreter: An LLM Agent For Data Science
- Autonomous Data Agents: A New Opportunity for Smart Data
- Data Agent: A Holistic Architecture for Orchestrating Data+AI Ecosystems
- GAIA: a benchmark for General AI Assistants
- DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
- KramaBench: A Benchmark for AI Systems on Data-to-Insight Pipelines over Data Lakes
- DABstep: Data Agent Benchmark for Multi-step Reasoning
- DSAEval: Evaluating Data Science Agents on a Wide Range of Real-World Data Science Problems
- Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
- Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- A Survey on LLM-as-a-Judge
- Agent-as-a-Judge: Evaluate Agents with Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection