A Statistical Audit of Physical AI Benchmark Redundancy
cs.RO, cs.AI
Submitted: 2026-08-26
Updated: 2026-08-29
Terminology
Sources
- RoboBrain 2.0 Technical Report
- Revealing the structure of language model capabilities
- Are We on the Right Way for Evaluating Large Vision-Language Models?
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- VisionArena: 230K Real World User-VLM Conversations with Preference Labels
- EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- BLINK: Multimodal Large Language Models Can See but Not Perceive
- Gemini Robotics 1.5: Pushing the Frontier of Generalist Robots with Advanced Embodied Reasoning, Thinking, and Motion Transfer
- Gemini Robotics: Bringing AI into the Physical World
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- DocVQA: A Dataset for VQA on Document Images
- Qwen3-VL Technical Report
- SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
- MindCube: Spatial Mental Modeling from Limited Views
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving