OnlineWM: Causality-Aware Active Online Learning for Effective World Modeling
cs.CV, cs.LG
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 23 pages, 9 figures. Project page: https://onlinewm.github.io/
Project page: https://onlinewm.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- MineStudio: A Streamlined Package for Minecraft AI Agent Development
- MineRL: A Large-Scale Dataset of Minecraft Demonstrations
- World Models
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- RELIC: Interactive Video World Model with Long-Horizon Memory
- Open-Endedness is Essential for Artificial Superhuman Intelligence
- Auto-Encoding Variational Bayes
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- Cameras as Relative Positional Encoding
- Open-Sora Plan: Open-Source Large Video Generation Model
- Depth Anything 3: Recovering the Visual Space from Any Views
- Flow Matching for Generative Modeling
- Muon is Scalable for LLM Training
- WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting
- Yume-1.5: A Text-Controlled Interactive World Generation Model
- WorldCam: Interactive Autoregressive 3D Gaming Worlds with Camera Pose as a Unifying Geometric Representation
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models