Code Plans, Diffusion Renders: Open-Ended Generative World Modeling
cs.CV
Submitted: 2026-09-22
Updated: 2026-09-22
Terminology
Sources
- AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
- MASS: Multiplayer World Models with Authoritative Shared State
- ShadowDancer: Teaching Video World Models Any Action by Learning Unified Dynamics Representations from a Video and Its Shadow
- Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
- DeepVerse: 4D Autoregressive Video Generation as a World Model
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- DreamX-World 1.0: A General-Purpose Interactive World Model
- LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
- Infinite Worlds with Versatile Interactions
- MineWorld: a Real-Time and Open-Source Interactive World Model on Minecraft
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Multiplayer Interactive World Models with Representation Autoencoders
- MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
- Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft
- ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
- Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models