Valerant: An Automatic Navigable Game Map Generator via Action-Conditioned World Model Exploration
cs.AI
Submitted: 2026-09-08
Updated: 2026-09-23
License: http://creativecommons.org/licenses/by/4.0/
The gist: World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior.
Terminology
Abstract
World Action Models (WAMs) couple predictive world modeling with action generation, allowing anticipated future states to guide agent behavior. Although WAMs are rapidly advancing embodied AI, general-purpose counterparts remain largely unexplored in games. Existing game-oriented approaches often combine action-conditioned world models with external policies and reward functions to realize WAM-like decision-making, yet they operate mainly in 2D visual observation space and do not instantiate persistent 3D geometry. Extending this paradigm to 3D games introduces a distinct challenge. In autonomous driving and robotics, the physical environment exists independently of the model, providing a persistent 3D world in which selected actions can be executed. Games have no such external substrate; the virtual world itself must be instantiated. Most playable games require a persistent and navigable space, while 3D games additionally require explicit geometry that supports movement and interaction. Action-conditioned video rollouts provide visual observations but not this spatial representation. We present Valerant, a training-free framework that transforms a pretrained action-conditioned world model into a WAM for exploring and constructing 3D game maps. By coupling predictive visual rollouts with SLAM-based spatial reconstruction and exploration-driven action selection, Valerant progressively transforms a single image into a persistent 3D game map. This framework extends WAM-based interaction beyond 2D visual simulation and offers a new approach to reducing manual effort in 3D game-map creation.
Sources
- COMBAT: Conditional World Models for Behavioral Agent Training
- Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control
- Learning to Explore using Active Neural SLAM
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- Beyond Pixel Histories: World Models with Persistent 3D State
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- World Models
- Dream to Control: Learning Behaviors by Latent Imagination
- World Reconstruction From Inconsistent Views
- HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- Evaluating Real-World Robot Manipulation Policies in Simulation
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- WorldSimBench: Towards Video Generation Models as World Simulators
- AVID: Adapting Video Diffusion Models to World Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- World Action Models: The Next Frontier in Embodied AI
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection