What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents
cs.AI, cs.CL
Submitted: 2026-08-27
Updated: 2026-08-27
Code: https://github.com/R2E-Gym/R2E-Gym
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: LLM agents increasingly rely on generated interaction data to learn how to interact with external environments.
Terminology
Abstract
LLM agents increasingly rely on generated interaction data to learn how to interact with external environments. Agentic data generation must maintain consistency among environments, tasks, interactions, and success signals while producing experience that is useful rather than merely abundant. Existing work spans many agent domains, but domain-centered organization and heterogeneous evaluation often obscure common generation mechanisms and conflate candidate construction with verification and selection. This work develops a two-level framework for the field. First, we represent agentic data as a common factorized object (E,q,τ,v), comprising an environment specification, task signal, interaction realization, and optional verifier. We organize generation paradigms by their primary anchor and dependency structure. Second, we formulate generation as constrained distribution design through the Accuracy-Complexity-divErsity (ACE) lens. Accuracy establishes the feasible support of grounded and internally consistent data. Within this support, Complexity places learning mass relative to the capability of a declared learner and execution configuration, while divErsity controls coverage and redundancy of data. Using this framework, we explore how prior work verifies generated experience, constructs and calibrates difficulty, and expands behavioral coverage. The literature reveals a shift toward execution-grounded accuracy, learner-relative complexity, and diversity beyond surface variation or dataset size. We further discuss broader directions and emerging trends in agentic data generation through the ACE lens, including their implications for scaling, data sources, training regimes and adaptive learning. Overall, the central challenge is not simply to generate more data, but to continually allocate valid, informative, and non-redundant experience as agents and environments evolve.
Sources
- Tool-R0: Self-Evolving LLM Agents for Tool-Learning from Zero Data
- Open-SWE-Traces: Advancing Dual-Mode Multilingual Distillation for Software Engineering Agents
- SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- AutoForge: Automated Environment Synthesis for Agentic Reinforcement Learning
- SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
- Safe and Scalable Web Agent Learning via Recreated Websites
- DIVE: Scaling Diversity in Agentic Task Synthesis for Generalizable Tool Use
- SWE-Universe: Scale Real-World Verifiable Environments to Millions
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- AgentFrontier: Expanding the Capability Frontier of LLM Agents with ZPD-Guided Data Synthesis
- Graph2Eval: Automatic Multimodal Task Generation for Agents via Knowledge Graphs
- OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis
- Simulating Complex Multi-Turn Tool Calling Interactions in Stateless Execution Environments
- Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence
- Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
- Toward Scalable Terminal Task Synthesis via Skill Graphs
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Self-Play Meets Skill Evolution: Self-Evolving Search Agents that Pose, Solve, and Remember
- From Self-Evolving Synthetic Data to Verifiable-Reward RL: Post-Training Multi-turn Interactive Tool-Using Agents
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection