DELE-w0.5: Inferring Action from Future Latent State for Robotic Manipulation
cs.RO, cs.AI, cs.CV, cs.LG
Submitted: 2026-08-22
Updated: 2026-08-31
Comments: DeepLeap Technology Co., Ltd., Shenzhen, China
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions.
Terminology
Abstract
World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling. For robotic manipulation, the goal of a world model is not to reproduce how the world looks at every intermediate moment, but to predict the state that the world will reach after an action is executed. The intermediate frames only describe the visual transition between physical states, which consumes substantial model capacity and computation, but do not directly specify the physical outcome that the robot action is intended to produce. In this paper, we propose DELE-w0.5, which infers robot actions from predicted future states without relying on video generation. Concretely, DELE-w0.5 infers the action sequence from its corresponding compact future latent state. The future latent state captures the action-relevant physical outcome of robot interaction and serves as an explicit bridge between world modeling and action generation. The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame. This formulation removes the high-dimensional visual redundancy introduced by dense video representations, and it therefore enables cheaper training and low-latency inference. Across 640 real-robot trials on four long-horizon manipulation tasks, our DELE-w0.5 achieves the best performance among all compared policies, attaining 62.5% overall full-task success and 81.3% macro ordered-stage progress. It outperforms the strongest baseline by 32.5 percentage points in full-task success and 20.1 percentage points in macro progress.
Sources
- Flash-WAM: Modality-Aware Distillation for World Action Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- Genie: Generative Interactive Environments
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- WorldVLA: Towards Autoregressive Action World Model
- GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- Light-WAM: Efficient World Action Models with State-Fusion Action Decoding
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving