IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models
cs.RO, cs.AI
Submitted: 2026-09-22
Updated: 2026-10-08
Comments: preprint
Code: https://github.com/xiaoqi-7/IndustrialVLA-Bench
License: http://creativecommons.org/licenses/by/4.0/
The gist: Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate
Terminology
Abstract
Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks and represent alternative design choices, they are commonly reported under different evaluation protocols, leaving their capability, robustness, language sensitivity, and deployment-cost trade-offs unclear. We present IndustrialVLA-Bench, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema. It separately evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost. Reported task scores aggregate three complete evaluations with distinct random seeds under a fixed checkpoint and inference configuration. Across all six systems, clean LIBERO averages differ by only 1.58 points, whereas robustness and paraphrase summaries span 14.62 and 31.08 points. Restricting every comparison to the three protocol-faithful systems preserves the effect (1.36, 14.62 and 23.10 points), so the diagnostic separation reported here does not depend on the weaker evidence tiers. We additionally report observed inference latency, peak memory, runtime mode, and an evidence status for every system. Protocol-faithful, near-reproduction, and pending-verification entries remain visibly separated; only protocol-faithful entries support strict comparisons. Rather than claiming universal superiority of either paradigm, IndustrialVLA-Bench provides traceable evidence for comparing released robot policies on shared practical criteria. Code and evaluation records are available at https://github.com/xiaoqi-7/IndustrialVLA-Bench.
Sources
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control and Planning
- RT-1: Robotics Transformer for Real-World Control at Scale
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- Xiaomi-Robotics-0: An Open-Sourced Vision-Language-Action Model with Real-Time Execution
- LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning
- VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models
- An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
- LeRobot: An Open-Source Library for End-to-End Robot Learning
- Do World Action Models Generalize Better than VLAs? A Robustness Study
- RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
- CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- LangGap: Diagnosing and Closing the Language Gap in Vision-Language-Action Models
- HazardArena: Evaluating Semantic Safety in Vision-Language-Action Models
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving