Exploring Autonomous Agentic Data Engineering for Model Specialization
cs.CL, cs.AI, cs.IR, cs.LG
Submitted: 2026-05-28
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026 main conference
Code: https://github.com/zjunlp/DataAgent
Project page: https://datapreparationbench.github.io
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data.
Terminology
Abstract
Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data. Existing LLM-based data curation methods primarily rely on human-designed workflows, leaving it unexamined whether LLMs can autonomously execute an end-to-end data engineering pipeline for model specialization. We formalize Autonomous Agentic Data Engineering, a novel task designed to evaluate LLMs as autonomous data engineers that drive model specialization through end-to-end data curation. We frame data as an optimizable component and study agents that plan, generate, and iteratively optimize training data across multiple domains, guided by post-training performance improvement. Experiments show that autonomous LLM data engineers yield substantial gains, as GPT-5.2 constructs a training curriculum that improves a student model by 57.29%, entirely through iterative, agent-driven data adaptation. By illuminating both potential and bottlenecks, our study establishes autonomous data engineering as a measurable capability and charts a path toward agent-driven model specialization (Code will be released at https://github.com/zjunlp/DataAgent).
Sources
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- FLEX: Continuous Agent Evolution via Forward Learning from Experience
- Evaluating Large Language Models Trained on Code
- MemPrivacy: Privacy-Preserving Personalized Memory Management for Edge-Cloud Agents
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Search, Verify and Feedback: Towards Next Generation Post-training Paradigm of Foundation Models via Verifier Engineering
- OpenThoughts: Data Recipes for Reasoning Models
- Textbooks Are All You Need
- AIDE: AI-Driven Exploration in the Space of Code
- Towards Active Synthetic Data Generation for Finetuning Language Models
- TACO: Topics in Algorithmic COde generation dataset
- DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI
- SkillOS: Learning Skill Curation for Self-Evolving Agents
- PostTrainBench: Can LLM Agents Automate LLM Post-Training?
- Large Language Models as Optimizers
- FinGPT: Open-Source Financial Large Language Models
- AutoSkill: Experience-Driven Lifelong Learning via Skill Self-Evolution
- Diversify and Conquer: Diversity-Centric Data Selection with Iterative Refinement
- Dr. Zero: Self-Evolving Search Agents without Training Data
- MMSkills: Towards Multimodal Skills for General Visual Agents
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering