PAN: A World Model for General, Actionable, and Long-Horizon World Simulation
cs.CV, cs.AI, cs.CL, cs.LG
Submitted: 2025-11-12
Updated: 2026-09-05
Code: https://github.com/nvidia-cosmos/cosmos-predict2
License: http://creativecommons.org/licenses/by/4.0/
The gist: A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whether spontaneously or in response to their actions, and
Terminology
Abstract
A world model is a cognitive simulator of the real-world environment allowing biological agents to reason about how the world evolves, whether spontaneously or in response to their actions, and accordingly to plan and strategize. In building Artificial Intelligence (AI) systems, world models represent the next frontier beyond large language models (LLMs) to enable physical and embodied intelligence in AI agents, allowing them to perform decision-making through simulative reasoning and reinforcement-learning through simulative trials. Recent advancements in world modeling have yielded impressive progress in video generation, 3-D scene evolution, robotic dynamics, and game simulation, but limitations persist in general, open-domain, action-driven prediction, long-horizon consistency, and abstract reasoning and planning. Moreover, fundamental architectural questions, whether it be state representation, information flow, or training objectives, remain unresolved. In this paper, we introduce PAN, a world model built on the Generative Latent Prediction (GLP) architecture. GLP combines stateful latent representations of world states; an encoder--decoder closed-loop information flow; an LLM/diffusion-based mixed reasoning backbone; and a non-degenerate generative reconstruction objective whose fidelity is ``dampable'' to balance fine-grained detail against semantic saliency. Compared to several existing systems, PAN demonstrates advantages beyond standard video generation in action-conditioned world simulation, long-horizon forecasting, and simulative reasoning and planning, capabilities we argue should serve as the primary criteria for evaluating world models.
Sources
- Cosmos World Foundation Model Platform for Physical AI
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Qwen2.5-VL Technical Report
- Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning
- UniMax: Fairer and more Effective Language Sampling for Large-Scale Multilingual Pretraining
- Flex Attention: A Programming Model for Generating Optimized Attention Kernels
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- WorldScore: A Unified Evaluation Benchmark for World Generation
- The Matrix: Infinite-Horizon World Generation with Real-Time Moving Control
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- World Models
- Dream to Control: Learning Behaviors by Latent Imagination
- Classifier-Free Diffusion Guidance
- Imagen Video: High Definition Video Generation with Diffusion Models
- GAIA-1: A Generative World Model for Autonomous Driving
- DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
- Model-Based Reinforcement Learning for Atari
- HunyuanVideo: A Systematic Framework For Large Video Generative Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models