CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model
cs.CV
Submitted: 2026-09-19
Updated: 2026-09-22
Code: https://github.com/AetherLabsAI/CausalWM
Project page: https://fysics-ai.github.io/Fysiverse-3D-project-page
Terminology
Sources
- Cosmos World Foundation Model Platform for Physical AI
- ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
- Qwen3-VL Technical Report
- CoPhy: Counterfactual Learning of Physical Dynamics
- RT-1: Robotics Transformer for Real-World Control at Scale
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- BWM: A Low-Cost High-Fidelity World Simulator for Robot Learning
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- FlowWAM: Optical Flow as a Unified Action Representation for World Action Models
- ComPhy: Compositional Physical Reasoning of Objects and Events from Videos
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- World Models
- LTX-Video: Realtime Video Latent Diffusion
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- Dream to Control: Learning Behaviors by Latent Imagination
- Mastering Atari with Discrete World Models
- Mastering Diverse Domains through World Models
- Vid2World: Crafting Video Diffusion Models to Interactive World Models
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models