Simulate to Generalize: Scaling Stateful Supervision for API-calling Agents using LLM World Models
cs.AI
Submitted: 2026-07-18
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by/4.0/
The gist: Training agents that generalize to unseen, stateful environments requires a massive dataset of state-changing trajectories covering a vast and diverse set of APIs.
Terminology
Abstract
Training agents that generalize to unseen, stateful environments requires a massive dataset of state-changing trajectories covering a vast and diverse set of APIs. However, scaling this broad supervision is severely bottlenecked by the immense effort required to implement and populate fully-executable environments across a broad spectrum of domains. To bypass this barrier, we introduce a data generation pipeline that decouples data synthesis from environment construction by leveraging LLMs as digital world models. Starting from only a list of broad domain names, our automated pipeline synthesizes diverse APIs and tasks. To produce trajectories, a teacher agent iteratively solves these tasks while an LLM simulator dynamically tracks state and provides coherent API responses on-the-fly. Finally, an automated judge filters the trajectories for quality. Fine-tuning on our broad synthetic dataset yields significant performance gains on AppWorld and OfficeBench, two challenging stateful benchmarks featuring environments completely unseen during training. These results establish our LLM world model-based synthesis approach as a highly scalable path for training generalizable, stateful API-calling agents.
Sources
- Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- GPT-4o System Card
- Simulating Environments with Reasoning Models for Agent Training
- Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
- ToolAlpaca: Generalized Tool Learning for Language Models with 3000 Simulated Cases
- OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
- Controllable and Verifiable Tool-Use Data Synthesis for Agentic Reinforcement Learning
- Qwen3 Technical Report
- GLM-5: from Vibe Coding to Agentic Engineering
- GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection