LM-X: Explainable Vision--Language--Action Modeling via Progress, Event, and Uncertainty Prediction
cs.RO, cs.LG
Submitted: 2026-08-26
Updated: 2026-10-08
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands.
Terminology
Abstract
Generalist vision--language--action (VLA) policies learn long-horizon behavior mainly through short-horizon action prediction and reveal little beyond sampled commands. This creates two coupled bottlenecks: a single action target must implicitly absorb task progress, intermediate intent, and local reliability, while these control states remain hidden during execution. Inspired by functional principles of biological sensorimotor control, we introduce LM-X, which organizes prediction across task, event, and motor scales without claiming anatomical correspondence. Three explicitly supervised signals are emitted online and directly condition action generation: return-to-go (RTG) measures visible task progress, event-to-go (ETG) identifies the next semantic transition, and heteroscedastic action flow estimates local reliability through propagated variance. Explanation is therefore intrinsic to control rather than generated post hoc. Before a costly 20-day pretraining run on 64 NVIDIA B200 GPUs, a controlled five-task pretraining gate verifies the design: the complete model improves success by 16.0 points over the action-only backbone and by 10.8 points over the strongest single-head variant. We then train LM-X on more than 20,000 hours of real-robot trajectories, including over 1,000 hours of failed policy rollouts. LM-X achieves 74.1% across 50 randomized-hard RoboTwin2.0 tasks versus 55.4% for GR00T N1.7, and 68.6% versus 50.7% across seven real-robot tasks. RTG tracks semantic progress and visible regression, while variance rises during hesitation and oscillatory control. These results show that explicit multi-timescale predictive state can strengthen control while exposing interpretable internal estimates.
Sources
- Cosmos 3: Omnimodal World Models for Physical AI
- Cosmos World Foundation Model Platform for Physical AI
- RT-H: Action Hierarchies Using Language
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- RT-1: Robotics Transformer for Real-World Control at Scale
- AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- SCALE: Self-uncertainty Conditioned Adaptive Looking and Execution for Vision-Language-Action Models
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- Flow Matching with Uncertainty Quantification and Guidance
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- Generative Uncertainty in Diffusion Models
- OpenVLA: An Open-Source Vision-Language-Action Model
- Why Does Action Chunking Improve Behavioral Cloning Performance in Robotic Control?
- Flow Matching for Generative Modeling
- Flow Matching Guide and Code
- STEAM: Self-Supervised Temporal Ensemble Advantage Modeling for Real-World Robot Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving