Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective
cs.LG, cs.CL
Submitted: 2025-10-16
Updated: 2026-09-09
Comments: Accepted to Findings of EMNLP 2026
Code: https://github.com/shiqichen17/SPA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large Language Models (LLMs) as agents often fail to improve in new environments.
Terminology
Abstract
Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call exploration collapse: under reinforcement learning (RL) in environments whose states are unfamiliar to the policy, Pass@k, the probability that at least one of k sampled trajectories succeeds, drops markedly over training even as Pass@1 edges up, revealing increasingly brittle exploration; environments closer to the pretraining distribution show no such decline. We trace this collapse to weak grounding in environment states and dynamics, and study a simple remedy: explicitly teaching the agent to estimate the current state and predict its transitions before optimizing for reward. We instantiate it as SPA, an explore-then-exploit recipe that cold-starts the policy with a Self-Experience supervised finetuning (SFT) stage, collecting the model's own interaction trajectories and supervising state and next-state prediction, and then runs standard RL. The resulting world model serves as a grounded initialization for RL rather than an inference-time planner. Across unseen environments, SPA consistently and substantially improves over vanilla RL: for example, it raises the Sokoban success rate from 25.6% to 59.8% on Qwen2.5-1.5B-Instruct, letting sub-3B models surpass a 20B baseline on these tasks. Controlled studies indicate that the gains track four factors: grounded state representations, explicit transition modeling, self-experience trajectories from a sufficiently strong exploration policy, and adequate coverage of transition data.
Sources
- GPT-4 Technical Report
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
- Proximal Policy Optimization Algorithms
- A Survey of Deep Reinforcement Learning in Video Games
- Large Language Models as Generalizable Policies for Embodied Tasks
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
- WebDancer: Towards Autonomous Information Seeking Agency
- Critique of World Model
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks