DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for an Efficient World-Action Model
cs.RO, cs.CV
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/deltawam/DeltaWAM
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Motus: A Unified Latent Action World Model
- A Short Note on the Kinetics-700 Human Action Dataset
- WorldVLA: Towards Autoregressive Action World Model
- LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies
- Foresight Without Seeing: Latent Futures for World Action Models
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- MolmoAct: Action Reasoning Models that can Reason in Space
- DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
- Causal World Modeling for Robot Control
- Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
- JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- LDA-1B: Scaling Latent Dynamics Action Model via Universal Embodied Data Ingestion
- mimic-video: Video-Action Models for Generalizable Robot Control Beyond VLAs
- DINOv3
- World Guidance: World Modeling in Condition Space for Action Generation
- LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving