LLM-BabyBench: Can Language Models Plan in Worlds They Can Simulate?
cs.AI, cs.CL
Submitted: 2025-05-17
Updated: 2026-09-11
License: http://creativecommons.org/licenses/by/4.0/
The gist: When an interactive benchmark reports a single success rate for a language-model agent, it is rarely clear what that number measures.
Terminology
Abstract
When an interactive benchmark reports a single success rate for a language-model agent, it is rarely clear what that number measures. A failure can come from perception, ambiguous instructions, retrieval, missing commonsense about what actions do, an incorrect model of the dynamics, or planning, and an aggregate score does not separate them. LLM-BabyBench recasts the procedurally generated BabyAI gridworld as a fully observable, purely textual environment in which every source of failure but planning is removed by construction. The whole grid is serialised into the prompt, instructions come from a small formal grammar, every object's coordinate is stated, the six actions and their effects are specified, and a deterministic expert validates each answer by executing it rather than judging it. On this substrate we define the PPD suite: Predict asks for the state that follows an action sequence, Plan for an action sequence that reaches a goal, and Decompose for a subgoal sequence that achieves a mission, scored by three assistance-aware metrics that separate understanding a mission from sequencing it. Across seven frontier and open models, simulation is far ahead of planning for every model, near saturation for the strongest and well short of it for the weakest, and the length of the required solution, not grid size or obstacle count, governs planning failure. Each model has a characteristic horizon beyond which single-attempt success collapses. Where enough instances are solved to support the ratio, returned plans stay near-optimal. Models that write out their working show why: they commit to one family of corridor-shaped route and verify it with no means of backing out, so what they return is either near-optimal or invalid. The same pattern holds one level up: decomposition precision falls to zero on long missions even where comprehension persists.
Sources
- On the Measure of Intelligence
- Training Verifiers to Solve Math Word Problems
- TextWorld: A Learning Environment for Text-based Games
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Imitation Learning from Suboptimal Demonstrations via Meta-Learning An Action Ranker
- The Llama 3 Herd of Models
- Inner Monologue: Embodied Reasoning through Planning with Language Models
- Do As I Can, Not As I Say: Grounding Language in Robotic Affordances
- BIG-Bench Extra Hard
- WikiHow: A Large Scale Text Summarization Dataset
- Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems
- LLMs Can Easily Learn to Reason from Demonstrations Structure, not content, is what matters!
- PLANET: A Collection of Benchmarks for Evaluating LLMs' Planning Capabilities
- 3DSRBench: A Comprehensive 3D Spatial Reasoning Benchmark
- GameTraversalBenchmark: Evaluating Planning Abilities Of Large Language Models Through Traversing 2D Game Maps
- GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs
- DocFinQA: A Long-Context Financial Reasoning Dataset
- Habitat: A Platform for Embodied AI Research
- ALFRED: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks
- ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection