vla-eval: A Unified Evaluation Harness for Vision-Language-Action Models
cs.AI
Submitted: 2026-03-14
Updated: 2026-09-14
Code: https://github.com/allenai/vla-evaluation-harness
Project page: https://allenai.github.io/vla-evaluation-harness/leaderboard
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- Memory, Benchmark & Robots: A Benchmark for Solving Complex Tasks with Reinforcement Learning
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- RoboMME: Benchmarking and Understanding Memory for Robotic Generalist Policies
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
- Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- FAST: Efficient Action Tokenization for Vision-Language-Action Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection