Devol-ONE: One Autoregressive Mixture of Transformers to Unify Vision-Language-Action and Latent World Modeling
cs.CV
Submitted: 2026-09-26
Updated: 2026-10-01
Project page: https://robotwin-platform.github.io/leaderboard
Terminology
Sources
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Motus: A Unified Latent Action World Model
- LAOF: Robust Latent Action Learning with Optical Flow Constraints
- AHA-WAM:Asynchronous Horizon-Adaptive World-Action Modeling with Observation-Guided Context Routing
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- IGOR: Image-GOal Representations are the Atomic Control Units for Foundation Models in Embodied AI
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- Learning Latent Action World Models In The Wild
- Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising
- Object-Centric Latent Action Learning
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
- Evo-0: Vision-Language-Action Model with Implicit Spatial Understanding
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Evo-Depth: A Lightweight Depth-Enhanced Vision-Language-Action Model
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- World Action Models: A Survey
- StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
- VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models