Causally Debiased Latent Action Model for Embodied Action-Conditioned World Models
cs.CV, cs.RO
Submitted: 2026-07-10
Updated: 2026-09-26
Terminology
Sources
- World Models
- Mastering Diverse Domains through World Models
- ACWM-Phys: Investigating Generalized Physical Interaction in Action-Conditioned Video World Models
- iVideoGPT: Interactive VideoGPTs are Scalable World Models
- Open X-Embodiment: Robotic Learning Datasets and RT-X Models
- Genie: Generative Interactive Environments
- Learning to Act without Actions
- DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos
- ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- SAM 3: Segment Anything with Concepts
- EgoDex: Learning Dexterous Manipulation from Large-Scale Egocentric Video
- CoWTracker: Tracking by Warping instead of Correlation
- Imitating Latent Policies from Observation
- Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos
- IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
- Learning Additively Compositional Latent Actions for Embodied AI
- Co-Evolving Latent Action World Models
- MVP-LAM: Learning Action-Centric Latent Action via Cross-Viewpoint Reconstruction
- Learning Interactive Real-World Simulators
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models