Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving
cs.CV, cs.CL, cs.RO
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: Accepted by EMNLP 2026
Code: https://github.com/tangzhengxu/awesome-av-cot
License: http://creativecommons.org/licenses/by/4.0/
The gist: Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer.
Terminology
Abstract
Chain-of-thought (CoT) reasoning powers generative models by eliciting intermediate steps before producing an answer. In autonomous driving, the answer is a continuous action. Thus its reasoning must share the same spatiotemporal structure as the physical world. This survey studies the resulting shift from textual CoT to action-grounded reasoning. Surveying 171 papers, including 130 method papers and 41 benchmarks, datasets, surveys, and analysis papers, we propose a representation-centered taxonomy that treats the form of the intermediate state as the organizing axis. We systematize the 130 methods into four categories: language-based, visual-spatial, latent-dynamic, and externalized reasoning, further divided into 13 subtypes tied to distinct regions of interests. Our synthesis shows that the open frontier of reasoning in driving agents lies in intermediate representations that can be grounded in the real world, coupled to real-time action, and verified under safety-critical systems. Project page: https://github.com/tangzhengxu/awesome-av-cot.
Sources
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models
- Bridging Large-Model Reasoning and Real-Time Control via Agentic Fast-Slow Planning
- V2V-LLM: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models
- Retrieval-Based Interleaved Visual Chain-of-Thought in Real-World Driving Scenarios
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- VERDI: VLM-Embedded Reasoning for Autonomous Driving
- AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
- MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
- Self-Supervised Bootstrapping of Action-Predictive Embodied Reasoning
- SteerVLA: Steering Vision-Language-Action Models in Long-Tail Driving Scenarios
- DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
- DiffVLA++: Bridging Cognitive Reasoning and End-to-End Driving through Metric-Guided Alignment
- RAD-LAD: Rule and Language Grounded Autonomous Driving in Real-Time
- Accelerating Structured Chain-of-Thought in Autonomous Vehicles
- Listen, Look, Drive: Coupling Audio Instructions for User-aware VLA-based Autonomous Driving
- AppleVLM: End-to-end Autonomous Driving with Advanced Perception and Planning-Enhanced Vision-Language Models
- DriveAction: A Benchmark for Exploring Human-like Driving Decisions in VLA Models
- AgentsCoDriver: Large Language Model Empowered Collaborative Driving with Lifelong Learning
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models