Generative AI for Autonomous Driving: Frontiers and Opportunities
cs.CV, cs.AI, cs.RO
Submitted: 2025-05-13
Updated: 2026-10-06
Code: https://github.com/taco-group/GenAI4AD
Terminology
Sources
- nuScenes: A multimodal dataset for autonomous driving
- Specification and Validation of Autonomous Driving Systems: A Multilevel Semantic Framework
- Automatic Generation of Scenarios for System-level Simulation-based Verification of Autonomous Driving Systems
- Moving Forward: A Review of Autonomous Driving Software and Hardware Systems
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Mcity Data Collection for Automated Vehicles Study
- A Comprehensive Survey of AI-Generated Content (AIGC): A History of Generative AI from GAN to ChatGPT
- LLaMA: Open and Efficient Foundation Language Models
- Code Llama: Open Foundation Models for Code
- Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- A Survey of World Models for Autonomous Driving
- Exploring the Interplay Between Video Generation and World Models in Autonomous Driving: A Survey
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- Generative AI in Transportation Planning: A Survey
- Vision-Language-Action (VLA) Models: Concepts, Progress, Applications and Challenges
- A Survey for Foundation Models in Autonomous Driving
- LLM4Drive: A Survey of Large Language Models for Autonomous Driving
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models