More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models

arXiv:2510.04532 · cs.AI, cs.CL, cs.RO · Submitted 2025-10-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models".

Tom: The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: Let's look at the setup here. The paper, "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models," introduces this whole DriveMind project to probe that causal relationship between reasoning and planning.

Jane: It’s interesting because they aren't just looking at whether a model can drive well; they're asking if the way it reasons translates into better driving decisions, or if it’s just using shortcuts.

Lu: The core idea is that if a model learns to exploit priors—like its current location or past movements—it might get high planning scores without actually needing the reasoning process to be causal.

Meng: So, the paper seems concerned that this reasoning might just be an accidental byproduct rather than something that directly leads to safe and correct trajectories.

Lalam: That makes sense from a practical standpoint; if the AI is relying on a shortcut, it might fail in an unexpected situation that requires true causal understanding.

The paper's summary: Tom: So, what’s the actual finding? The researchers found a consistent causal disconnect: when you remove those textual priors—the ego state and history—the planning scores take big drops, but removing the Chain-of-Thought process only causes minor changes.

Jane: That suggests that the training has yielded reasoning that's more like a byproduct than the actual cause of good planning performance. It’s not directly driving the plan.

Lu: The paper posits this as the Reasoning-Planning Decoupling Hypothesis, which suggests that what we train up as reasoning might just be an ancillary byproduct of learning these textual priors instead of being a causal mediator in itself.

Meng: If that's true, then current training methods might be insufficient because they aren't actually forging that direct link between thought and action.

Lalam: It means we need a way to test this causality rigorously, which is why they built DriveMind to create the necessary data for this investigation into the paper "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models."

The paper's improvements: Tom: To diagnose this disconnect, they introduced two diagnostic tools. First, there’s a training-free probe called the Causal Probe that checks planning robustness against small input changes to see how much the agent relies on those textual priors.

Jane: And then they have lateral direction inversion, which tests if what the agent *says* it's reasoning about matches what it’s actually doing in its planning logic.

Lu: The results from these tools were quite telling: during the planning phase, attention to textual priors skyrockets from eleven point five two percent up to nearly twenty-seven percent, while attention to image tokens drops significantly, down below two percent.

Meng: That tells us that even when the AI is trying to plan a move, it’s heavily focused on its history and state information instead of looking at what's happening visually right now.

Lalam: So the authors are showing that while reasoning might be present during generation, it’s not influencing the actual planning step much, which points toward where we need to focus our improvements.

Conclusion: Tom: So to wrap up this discussion on "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models," it seems like current training paradigms are leading agents to learn shortcuts from textual priors rather than building a true causal link.

Jane: The paper suggests that the reasoning generated is often a plausible byproduct, but not necessarily something that causes the plan to be safe or correct. This means we need new training strategies that force the connection between thought and action.

Lu: Their future work focuses on two main directions: first, mitigating modality bias through contrastive pre-finetuning to make vision essential, and second, breaking those shortcuts with contrastive learning using negative examples to penalize reliance on 'prior equals plan' shortcuts.

Meng: From an engineering standpoint, making vision indispensable is a big challenge because you have to design the training to enforce that visual input is the only thing that matters for the final decision.

Lalam: I think those contrastive learning ideas are where we can go next, trying to guide the policy away from those easy textual shortcuts and toward genuine causal understanding.

Tom: That's a lot of heavy lifting for future models, but it really highlights that we need causally-aware training methods if we want driving agents that are truly robust.

S-Lab, Nanyang Technological University

cs.AI, cs.CL, cs.RO

Submitted: 2025-10-06

Updated: 2026-10-08

Comments: NeurIPS 2026

Code: https://github.com/ApolloAuto/apollo

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings, traffic

Key concepts

Reasoning-Planning Disconnect Hypothesis
This hypothesis proposes that the reasoning an agent develops during training is just a byproduct, not the actual cause of its planning behavior. Agents trained only on text priors can perform well in planning without needing visual input or CoT reasoning, proving reasoning isn't always causal.
Causal Probe
A new diagnostic tool used to measure how much an agent relies on textual priors. It tests robustness by seeing if the agent's planning changes significantly when minor perturbations are made to these text-based inputs, revealing reliance on shortcuts.
Sequence-level Attention Analysis
This technique measures where an agent focuses its attention during different stages. The study found that while reasoning involves visual input, the planning phase heavily prioritizes textual priors over actual image tokens, showing a shift away from visual grounding during action selection.

Terminology

Summary

The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings, traffic signals) and the CoT reasoning

DriveMind Dataset Creation

The researchers built DriveMind to investigate the causal relationship between reasoning and planning in VLM-based driving agents The dataset was constructed based on nuPlan as a foundation, specifically curated to facilitate the causal analysis of VLM-based driving agents DriveMind covers approximately 50, 000 samples spanning 61 driving scenarios, providing broad and diverse coverage for analysis The data generation pipeline involves three main stages: scene parsing and feature extraction, GPT-4.1-based CoT generation, and human verification The first stage extracts ego-vehicle priors such as current and historical states, as well as the global navigation goal

Testing the Reasoning-Planning Disconnect

The core finding is a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes This leads to the ReasoningPlanning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator An agent trained solely on textual priors, with no visual input and no CoT reasoning, achieves planning scores that match or even exceed those of a fully multimodal counterpart

Diagnostic Tools for Causal Analysis

To enable efficient diagnosis, the authors introduced two key diagnostic tools First, they introduced a novel, training-free probe called the Causal Probe that measures an agent’s reliance on priors by evaluating its planning robustness against minor input perturbations The first method is lateral offset perturbation, which tests if a robust agent should be stable under minor perturbations to textual priors The second method is lateral direction inversion, which checks for a contradiction between the agent’s stated reasoning and its actual planning logic

Sequence-Level Attention Analysis

The authors employed Sequence-level Attention Analysis to interpret information flow by measuring how much attention the planning sequence pays to preceding CoT tokens versus other contextual information They found that when generating the Reasoning sequence, attention to Image Tokens progressively increases, indicating a healthy process where reasoning is grounded in visual input Conversely, during the Planning phase, attention to Textual Priors skyrockets from 11.52% to nearly 27%, while attention to the Image Tokens becomes negligible

Conclusion and Future Directions

The work concludes that current training paradigms are insufficient to forge a causal link between reasoning and planning, leading agents to instead learn shortcuts from textual priors Future work plans focus on two directions: Mitigating Modality Bias via Contrastive Pre-Finetuning, which aims to make the visual input indispensable and Breaking Shortcuts via Contrastive Learning, which uses negative examples to guide policy and penalize reliance on prior equals plan shortcuts The combination of these strategies promises a driving agent whose planning is grounded in its own reasoning

How it works

The DriveMind dataset generation pipeline involves extracting ego-vehicle priors such as current and historical states, as well as the global navigation goal For visual information, camera images are rescaled and stitched into a three-row image grid corresponding to the front, side, and rear view groups The second stage uses GPT-4.1 to generate high-quality CoT by grounding the reasoning process with structured scene ground truth Finally, the training set is constructed from 61 driving scenarios sourced from the nuPlan training split

Reasoning-Planning Disconnect in VLM Driving Agents

The authors systematically tested the impact of information ablation by removing key input modalities such as visual information, textual priors, or CoT reasoning They found that an agent trained with no visual input and no CoT reasoning achieves planning scores that match or even exceed those of a fully multimodal counterpart This reliance on shortcuts is so entrenched that even applying advanced policy alignment methods like Group Relative Policy Optimization (GRPO) fails to substantively restore the causal link from reasoning to planning

Shortcut Learning and Causal Probe Results

The quantitative test termed lateral offset perturbation shows that both agents exhibit extreme sensitivity to this dynamic perturbation, with final lateral deviation far exceeding a single lane width The qualitative probe using lateral direction inversion reveals a stark contradiction between the agent’s stated reasoning and its actual planning logic This provides definitive evidence for the Reasoning-Planning Decoupling Hypothesis, as the CoT grpo agent generates correct reasoning but produces a wildly divergent and unsafe trajectory

Attention Distribution Findings

The sequence-level attention analysis shows that when generating the Planning sequence, attention to Textual Priors skyrockets from 11.52% to nearly 27%, while attention to the Image Tokens becomes negligible, dropping below 2% This finding suggests that current VLM driving agents heavily rely on textual priors during planning, with reasoning processes being almost irrelevant

Final Summary

In summary, the authors provide a new dataset and a diagnostic tool to evaluate the causal fidelity of future models They have found that agents learn to shortcut, relying predominantly on textual priors for planning, while their generated CoT often serves as a plausible but non-causal byproduct This work has revealed that the perceived interpretability of current agents can be misleading, as the reasoning are not causally linked to the final action and has underscored the need for causally-aware training paradigms to build truly robust driving agents

References

The references cited in the paper include works by Chen et al. (2024), Hu et al. (2023), Jiang et al. (2023), Wen et al. (2024) and Sima et al. (2025) as well as Geirhos et al. (2020), Yuan et al. (2024) and Delavari et al. (2025) and Dosovitskiy et al. (2017) and Bubeck et al. (2023), Ahn et al. (2024), Jiang et al. (2025) and Dosovitskiy et al. (2021) and Liu et al. (2023b) and Wang et al. (2025), Bai et al. (2025), Liu et al. (2023a) and Wang et al. (2025) and Omnidrive Wang et al. (2025) as well as DeepSeek-AI et al. (2025), Elahe Delavari et al. (2025), and Motional (2023). The references are listed on Page 16, 17, 18, and 9.

Improvements for AI systems

  1. Bold header: Causal Fidelity Verification via DriveMind Dataset

This improvement involves utilizing DriveMind to rigorously test whether planning is causally driven by this reasoning, as stated in the abstract. The resulting AI system can now be evaluated not just on trajectory quality, but on the causal link between reasoning and planning, moving beyond assessing how well the planning appears.

  1. Bold header: Introduction of a Training-Free Causal Probe

Implement a novel, training-free diagnostic method to measure an agent’s reliance on priors by evaluating its planning robustness against minor input perturbations. This allows for a disproportionately large degradation in planning scores following perturbation as evidence of shortcut learning, providing a new tool for evaluating the causal fidelity of VLM agents.

  1. Bold header: Modality-Aware Contrastive Pre-Finetuning

Introduce a dedicated contrastive prefinetuning stage designed to make the visual input indispensable, forcing the model to extract decisive information from vision. This aims to mitigate modality bias, addressing the finding that an agent with no visual input (Plan NoV) can achieve planning scores nearly identical to a fully multimodal agent.

  1. Bold header: Contrastive Learning for Shortcut Breaking

Enhance SFT/GRPO processes via contrastive learning by augmenting training data with verified conflict samples. This forces the policy to abandon the 'prior equals plan' shortcut and instead learn to genuinely reason, aiming to connect the reasoning and planning.

Abstract

Vision-Language Model (VLM) driving agents promise explainable end-to-end autonomy by first producing natural-language reasoning and then predicting trajectory planning. However, whether planning is causally driven by this reasoning remains a critical but unverified assumption. To investigate this, we build DriveMind, a large-scale driving Visual Question Answering corpus with plan-aligned Chain-of-Thought (CoT), automatically generated from nuPlan. Our data generation process converts sensors and annotations into structured inputs and, crucially, separates priors from to-be-reasoned signals, enabling clean information ablations. Using DriveMind, we train representative VLM agents with Supervised Fine-Tuning and Group Relative Policy Optimization and evaluate them with nuPlan's metrics. Our results, unfortunately, indicate a consistent causal disconnect in reasoning-planning: removing ego/navigation priors causes large drops in planning scores, whereas removing CoT produces only minor changes. Attention analysis further shows that planning primarily focuses on priors rather than the CoT. Based on this evidence, we propose the Reasoning-Planning Decoupling Hypothesis, positing that the training-yielded reasoning is an ancillary byproduct rather than a causal mediator. To enable efficient diagnosis, we introduce a novel, training-free probe that measures an agent's reliance on priors by evaluating its planning robustness against minor input perturbations. In summary, we provide the community with a new dataset and a diagnostic tool to evaluate the causal fidelity of future models.

Sources

Related papers