WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
cs.CV, cs.AI
Submitted: 2025-11-25
Updated: 2026-08-31
Code: https://github.com/sjyu001/WaymoQA
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their application to autonomous driving.
Terminology
Abstract
Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their application to autonomous driving. However, high-level reasoning in safety-critical scenarios, where avoiding one traffic risk can create another, remains a major challenge. Such reasoning is often infeasible with only a single front view and requires a comprehensive view of the environment, which we achieve through multi-view inputs. We define Safety-Critical Reasoning as a new task that leverages multi-view inputs to address this challenge. Then, we distill Safety-Critical Reasoning into two stages: first resolve the immediate risk, then mitigate the decision-induced downstream risks. To support this, we introduce WaymoQA, a dataset of 35,000 human-annotated question-answer pairs covering complex, high-risk driving scenarios. The dataset includes multiple-choice and open-ended formats across both image and video modalities. Experiments reveal that existing MLLMs underperform in safety-critical scenarios compared to normal scenes, but fine-tuning with WaymoQA significantly improves their reasoning ability, highlighting the effectiveness of our dataset in developing safer and more reasoning-capable driving agents. Our code and data are provided in https://github.com/sjyu001/WaymoQA
Sources
- VDT-Auto: End-to-end Autonomous Driving with VLM-Guided Diffusion Transformers
- GAIA-1: A Generative World Model for Autonomous Driving
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- AutoDrive-QA: A Multiple-Choice Benchmark for Vision-Language Evaluation in Urban Autonomous Driving
- LLaVA-OneVision: Easy Visual Task Transfer
- AVD2: Accident Video Diffusion for Accident Video Description
- Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models
- Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
- Qwen2.5-VL Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- VERDI: VLM-Embedded Reasoning for Autonomous Driving
- Mini-InternVL: A Flexible-Transfer Pocket Multimodal Model with 5% Parameters and 90% Performance
- Ferret: Refer and Ground Anything Anywhere at Any Granularity
- Open-Sora: Democratizing Efficient Video Production for All
- Guided Conditional Diffusion for Controllable Traffic Simulation
- SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain
- OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models