Colosseum V2: Benchmarking Generalization for Vision-Language-Action Models
cs.RO
Submitted: 2026-05-26
Updated: 2026-09-29
Comments: Accepted to IEEE Robotics and Automation Letters (RA-L)
Code: https://github.com/huggingface/lerobot
Project page: https://colosseum-v2.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- SAM 2: Segment Anything in Images and Videos
- robosuite: A Modular Simulation Framework and Benchmark for Robot Learning
- VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks
- VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation
- LIBERO-Para: A Diagnostic Benchmark and Metrics for Paraphrase Robustness in VLA Models
- ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
- VIP: Towards Universal Visual Reward and Representation via Value-Implicit Pre-Training
- PaLM-E: An Embodied Multimodal Language Model
- RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
- OpenVLA: An Open-Source Vision-Language-Action Model
- Provable Preconditioned Plug-and-Play Approach for Compressed Sensing MRI Reconstruction
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- Learning Transferable Visual Models From Natural Language Supervision
- Sigmoid Loss for Language Image Pre-Training
- MolmoAct: Action Reasoning Models that can Reason in Space
- VIMA: General Robot Manipulation with Multimodal Prompts
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving