DeliveryGym: An RL Environment for Long-Horizon Embodied Agent Planning with Adaptive Curriculum
cs.LG
Submitted: 2026-09-17
Updated: 2026-09-23
Code: https://github.com/MoonshotAI/Kimi-K3
License: http://creativecommons.org/licenses/by/4.0/
The gist: Executable environments enable LLM agents to learn from the consequences of their actions.
Terminology
Abstract
Executable environments enable LLM agents to learn from the consequences of their actions. For embodied agents, those consequences extend beyond whether the current task succeeds: completing a delivery can consume the time, energy, or money needed for later work. Learning to plan therefore requires environments that preserve these dependencies and turn them into feedback across a complete trajectory. We introduce DeliveryGym, a 3D environment for evaluating and training agents on continuous courier shifts. It couples multimodal tool interaction with persistent world dynamics and computes trajectory rewards from simulator events, making the costs of an agent's decisions available for reinforcement learning (RL). The environment also adapts future training shifts to the policy's observed weaknesses while keeping evaluation fixed. Across six models and 13 city maps, evaluation exposes a gap between reliably executing assigned deliveries and choosing and sequencing work over a shift. On the fixed test suite, RL improves Qwen3-VL-4B's net income by 54.3%, showing that learning from complete shifts improves performance under these coupled constraints. Adapting the training environment improves test income by 16.5% over uniform sampling at the same rollout budget, indicating that which situations an agent practices also matters. DeliveryGym provides an executable setting for studying how agents learn to coordinate deliveries and preserve resources for later orders within an episode.
Sources
- Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents
- Qwen3-VL Technical Report
- From Trainee to Trainer: LLM-Designed Training Environment for RL with Multi-Agent Reasoning
- AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
- EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
- EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies
- MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
- Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
- VerlTool: Towards Holistic Agentic Reinforcement Learning with Tool Use
- SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
- SimWorld Studio: Automatic Environment Generation with Evolving Coding Agent for Embodied Agent Learning
- The Interplay of Harness Design and Post-Training in LLM Agents
- AI2-THOR: An Interactive 3D Environment for Visual AI
- BEHAVIOR-1K: A Human-Centered, Embodied AI Benchmark with 1,000 Everyday Activities and Realistic Simulation
- OmniParser for Pure Vision Based GUI Agent
- DeliveryBench: Can Agents Earn Profit in Real World?
- MemGPT: Towards LLMs as Operating Systems
- Gorilla: Large Language Model Connected with Massive APIs
- SimWorld: An Open-ended Realistic Simulator for Autonomous Agents in Physical and Social Worlds
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks