StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
cs.AI
Submitted: 2025-07-10
Updated: 2026-08-26
Comments: Accepted by ECCV 2026. Project website: https://weihaotan.github.io/StarDojo
Project page: https://weihaotan.github.io/StarDojo
License: http://creativecommons.org/licenses/by/4.0/
The gist: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously.
Terminology
Abstract
Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate these skills simultaneously. To bridge this gap, we introduce StarDojo, a novel benchmark based on Stardew Valley, designed to assess AI agents in open-ended production-living simulations. In StarDojo, agents are tasked to perform essential livelihood activities such as farming and crafting, while simultaneously engaging in social interactions to establish relationships within a vibrant community. StarDojo features 1,000 meticulously curated tasks across five key domains: farming, crafting, exploration, combat, and social interactions. Additionally, we provide a compact subset of 100 representative tasks for efficient model evaluation. The benchmark offers a unified, user-friendly interface that eliminates the need for keyboard and mouse control, supports all major operating systems, and enables the parallel execution of multiple environment instances, making it particularly well-suited for evaluating the most capable foundation agents, powered by multimodal large language models (MLLMs). Extensive evaluations of state-of-the-art MLLMs agents demonstrate substantial limitations, with the best-performing model, GPT-4.1, achieving only a 12.7% success rate, primarily due to challenges in visual understanding, multimodal reasoning and low-level manipulation. As a user-friendly environment and benchmark, StarDojo aims to facilitate further research towards robust, open-ended agents in complex production-living environments.
Sources
- GPT-4 Technical Report
- WOLF: Werewolf-based Observations for LLM Deception and Falsehoods
- Qwen2.5-VL Technical Report
- Matterport3D: Learning from RGB-D Data in Indoor Environments
- MineRL: A Large-Scale Dataset of Minecraft Demonstrations
- Benchmarking the Spectrum of Agent Capabilities
- MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework
- VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
- AI2-THOR: An Interactive 3D Environment for Visual AI
- iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks
- AvalonBench: Evaluating LLMs Playing the Game of Avalon
- Continuous control with deep reinforcement learning
- BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games
- Habitat 3.0: A Co-Habitat for Humans, Avatars and Robots
- CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making Agents
- The StarCraft Multi-Agent Challenge
- Proximal Policy Optimization Algorithms
- Gemma 3 Technical Report
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection