Robusto-2: Benchmarking Humans & VLMs for Autonomous Driving in Lima & New York City
cs.CV, cs.AI, cs.RO
Submitted: 2026-06-18
Updated: 2026-09-12
Comments: Accepted as a Workshop paper at the Drive-X workshop (Foundation Models for Autonomous Driving) at the European Conference in Computer Vision (ECCV), 2026. Malmo, Sweden. 17 pages main body. 51 pages total. Data publicly available online
License: http://creativecommons.org/licenses/by/4.0/
The gist: As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new
Terminology
Abstract
As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action models; how well will these systems generalize in new settings, in particular out-of-distribution (OOD) edge-case scenarios in new geographies? In this paper, we study this open question by providing a full factorial analysis with human drivers of Lima, human drivers from New York City, and VLMs and showing them dashcam footage collected from Lima and New York City -- prompting them with a variety of questions under a Visual Question Answering (VQA) paradigm. In particular, we pick these two cities as they are highly challenging driving locations where no Self-Driving Car company currently operates in, and ask questions that span 4 categories: Factual, Ratings, Counterfactual and Reasoning. We find that Humans and VLMs diverge in their responses -- though this is modulated by the type of questions asked, and that Humans answer similarly independent of where they are from (Lima/NYC). To our surprise, we did not find a strong difference in terms of answers (Humans or VLMs) that was modulated by geography, likely due to their high out-of-distribution nature. Our dataset is available at: https://huggingface.co/datasets/Artificio/robusto-2
Sources
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- MIRAGE: The Illusion of Visual Understanding
- Qwen3-VL Technical Report
- Joint rotational invariance and adversarial training of a dual-stream Transformer yields state of the art Brain-Score for Area V4
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks
- SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
- Fail2Drive: Benchmarking Closed-Loop Driving Generalization
- Actions as Language: Fine-Tuning VLMs into VLAs Without Catastrophic Forgetting
- The Platonic Representation Hypothesis
- RealEngine: Simulating Autonomous Driving in Realistic Context
- OpenVLA: An Open-Source Vision-Language-Action Model
- Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
- Stress and Adaptation: Applying Anna Karenina Principle in Deep Learning for Image Classification
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- Uncovering Bias in Foundation Models: Impact, Testing, Harm, and Mitigation
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- MiniCPM-V: A GPT-4V Level MLLM on Your Phone
- VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models