WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory
cs.CV, cs.AI, cs.GR
Submitted: 2026-09-21
Updated: 2026-09-21
Comments: Project webpage: https://drexubery.github.io/WorldCrafter
Project page: https://drexubery.github.io/WorldCrafter
License: http://creativecommons.org/licenses/by/4.0/
The gist: Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints.
Terminology
Abstract
Video world models enable interactive exploration of dynamic environments, yet struggle to respect prior observations over long horizons and across viewpoints. We present WorldCrafter, a video world model that learns a camera-queryable implicit 3D-aware memory for this purpose. The key insight is to let the requested viewpoint shape how multi-view evidence is compressed into the video generator's limited token budget. Trained jointly with the video generator, a memory encoder and pose-conditioned readout module integrate historical observations into a fixed set of target view-specific tokens before denoising, without explicit depth-based correspondences. By combining this memory with recent temporal context and few-step distillation, WorldCrafter enables streaming scene exploration from a single input image or text prompt. Experiments across static and dynamic scenes show substantial gains in long-horizon consistency and camera-control accuracy while preserving visual quality during minute-scale exploration.
Sources
- AlayaWorld: Long-Horizon and Playable Video World Generation
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- Qwen2.5-VL Technical Report
- SkyReels-V2: Infinite-length Film Generative Model
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation
- DreamX-World 1.0: A General-Purpose Interactive World Model
- Infinite Worlds with Versatile Interactions
- Cosmos 3: Omnimodal World Models for Physical AI
- Compression and Retrieval: Implicit Memory Retrieval for Video World Models
- MAGI-1: Autoregressive Video Generation at Scale
- Lyra 2.0: Explorable Generative 3D Worlds
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- I3DM: Implicit 3D-aware Memory Retrieval and Injection for Consistent Video Scene Generation
- Open-Sora Plan: Open-Source Large Video Generation Model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models