Astronex-World 1.0: Real-Time Interactive World Model Foundation
cs.CV, cs.AI, cs.RO
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: Technical report. 25 pages, 13 figures, 10 tables. Project page: https://world.astronex.com.cn ; Code: https://github.com/Astronex-Robotics/Astronex-World ; Weights: https://huggingface.co/Astronex-Lab/Astronex-World
Code: https://github.com/Astronex-Robotics/Astronex-World
Project page: https://world.astronex.com.cn
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: We present Astronex-World 1.0, an open controllable video world-model foundation.
Terminology
Abstract
We present Astronex-World 1.0, an open controllable video world-model foundation. Given a text prompt (text-to-video) or an initial observation (image-to-video), the model predicts future visual states under frame-aligned camera trajectories, continuous actions, and an embodiment identifier, and accepts text events inserted at a specified position of a rollout. The family provides a bidirectional model for full-context generation and a causal model with block-causal attention and cross-block KV caching for persistent generation, both built on the Wan2.2-TI2V-5B prior. PRoPE injects camera intrinsics and extrinsics, while a 64-dimensional action stream modulates every Transformer layer. A five-stage training path develops bidirectional camera and action control, converts the backbone to block-causal generation, distills a few-step student, restores mixed-domain dynamics, and applies asymmetric DMD/DMD2 distribution matching. The causal model generates 832x480 video at 24 fps. All five training stages run on two NVIDIA L20 48 GB GPUs, and the causal model streams in real time on one. It scores 73.5 on WBench Navi and 70.0 on WBench Full. On Full, this 5B model is above the 13.6B LongCat-Video and the 14B Helios, within one point of the 22B LTX-2.3, and above YUME 1.5, which is post-trained from the same 5B prior on NVIDIA A100 GPUs. The reserved action input and output interfaces allow post-training for embodied intelligence and autonomous driving.
Sources
- Wan: Open and Advanced Large-Scale Video Generative Models
- World Models
- GAIA-1: A Generative World Model for Autonomous Driving
- Cosmos World Foundation Model Platform for Physical AI
- Mastering Diverse Domains through World Models
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- Yume: An Interactive World Generation Model
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- Advancing Open-source World Models
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
- MAGI-1: Autoregressive Video Generation at Scale
- SkyReels-V2: Infinite-length Film Generative Model
- Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
- Context as Memory: Scene-Consistent Interactive Long Video Generation with Memory Retrieval
- WorldMem: Long-term Consistent World Simulation with Memory
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models