Can Generalist Agents Automate Data Curation?
cs.AI, cs.CL, cs.CV, cs.ET, cs.LG
Submitted: 2026-06-02
Updated: 2026-09-19
Comments: Published as a Main Conference paper at EMNLP 2026
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against
Terminology
Abstract
Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against noisy benchmark feedback. We ask whether generalist coding agents can automate this data-curation loop. We introduce *Curation-Bench*, an agent-centric benchmark that fixes the model, training recipe, and evaluation suite while giving agents command-line access to inspect data, implement policies, submit them to a fixed training/evaluation pipeline, and revise. In a vision-language instruction-tuning instantiation, out-of-the-box agents reach strong published data-selection baselines within ten iterations. However, trajectory analysis reveals a persistent *execution-research gap*: agents mainly tune local policy variants rather than explore new policy families, even when given strategy guides and paper references. Scaffolds requiring each iteration to cite, instantiate, and adapt a prior method shift agents toward method-guided exploration. The scaffolded agent autonomously composes -- without human design input -- a data-selection policy that outperforms strong published baselines at one-tenth their data budget. Overall, current agents can run the curation loop, but reliable data research requires scaffolded method adaptation, not open-ended prompting alone. Code and benchmark are open-sourced.
Sources
- SemDeDup: Data-efficient learning at web-scale through semantic deduplication
- Qwen2.5-VL Technical Report
- MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
- Olmix: A Framework for Data Mixing Throughout LM Development
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
- AgentBench: Evaluating LLMs as Agents
- AdaDeDup: Adaptive Hybrid Data Pruning for Efficient Large-Scale Object Detection Training
- Data-efficient pre-training by scaling synthetic megadocs
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- SmolVLM: Redefining small and efficient multimodal models
- Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Investigating the Scaling Effect of Instruction Templates for Training Multimodal Language Model
- ICONS: Influence Consensus for Vision-Language Data Selection
- LESS: Selecting Influential Data for Targeted Instruction Tuning
- The AI Scientist-v2: Workshop-Level Automated Scientific Discovery via Agentic Tree Search
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection