More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
summary
The gist
The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings, traffic
In short
Researchers created a dataset to test if reasoning and planning are causally linked in Vision-Language Models for driving. They found a major disconnect: agents learn shortcuts from text priors rather than using their reasoning process for planning. This suggests current training methods create plausible but non-causal reasoning, requiring new training approaches.
Key concepts
- Reasoning-Planning Disconnect Hypothesis
- This hypothesis proposes that the reasoning an agent develops during training is just a byproduct, not the actual cause of its planning behavior. Agents trained only on text priors can perform well in planning without needing visual input or CoT reasoning, proving reasoning isn't always causal.
- Causal Probe
- A new diagnostic tool used to measure how much an agent relies on textual priors. It tests robustness by seeing if the agent's planning changes significantly when minor perturbations are made to these text-based inputs, revealing reliance on shortcuts.
- Sequence-level Attention Analysis
- This technique measures where an agent focuses its attention during different stages. The study found that while reasoning involves visual input, the planning phase heavily prioritizes textual priors over actual image tokens, showing a shift away from visual grounding during action selection.
Terminology used across episodes
This episode discusses
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models · Paper Radio
- Qwen2.5-VL Technical Report
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator
- End-to-End Autonomous Driving without Costly Modularization and 3D Manual Annotation
- DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
The paper
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models · Read on arXiv
S-Lab, Nanyang Technological University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models".
Tom: The gist The planning module from an agent’s output predominantly relies on textual priors (i.e., ego state, history) as shortcuts, largely ignoring the visual context (i.e., surroundings,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Let's look at the setup here. The paper, "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models," introduces this whole DriveMind project to probe that causal relationship between reasoning and planning.
Jane: It’s interesting because they aren't just looking at whether a model can drive well; they're asking if the way it reasons translates into better driving decisions, or if it’s just using shortcuts.
Lu: The core idea is that if a model learns to exploit priors—like its current location or past movements—it might get high planning scores without actually needing the reasoning process to be causal.
Meng: So, the paper seems concerned that this reasoning might just be an accidental byproduct rather than something that directly leads to safe and correct trajectories.
Lalam: That makes sense from a practical standpoint; if the AI is relying on a shortcut, it might fail in an unexpected situation that requires true causal understanding.
The paper's summary: Tom: So, what’s the actual finding? The researchers found a consistent causal disconnect: when you remove those textual priors—the ego state and history—the planning scores take big drops, but removing the Chain-of-Thought process only causes minor changes.
Jane: That suggests that the training has yielded reasoning that's more like a byproduct than the actual cause of good planning performance. It’s not directly driving the plan.
Lu: The paper posits this as the Reasoning-Planning Decoupling Hypothesis, which suggests that what we train up as reasoning might just be an ancillary byproduct of learning these textual priors instead of being a causal mediator in itself.
Meng: If that's true, then current training methods might be insufficient because they aren't actually forging that direct link between thought and action.
Lalam: It means we need a way to test this causality rigorously, which is why they built DriveMind to create the necessary data for this investigation into the paper "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models."
The paper's improvements: Tom: To diagnose this disconnect, they introduced two diagnostic tools. First, there’s a training-free probe called the Causal Probe that checks planning robustness against small input changes to see how much the agent relies on those textual priors.
Jane: And then they have lateral direction inversion, which tests if what the agent *says* it's reasoning about matches what it’s actually doing in its planning logic.
Lu: The results from these tools were quite telling: during the planning phase, attention to textual priors skyrockets from eleven point five two percent up to nearly twenty-seven percent, while attention to image tokens drops significantly, down below two percent.
Meng: That tells us that even when the AI is trying to plan a move, it’s heavily focused on its history and state information instead of looking at what's happening visually right now.
Lalam: So the authors are showing that while reasoning might be present during generation, it’s not influencing the actual planning step much, which points toward where we need to focus our improvements.
Conclusion: Tom: So to wrap up this discussion on "More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models," it seems like current training paradigms are leading agents to learn shortcuts from textual priors rather than building a true causal link.
Jane: The paper suggests that the reasoning generated is often a plausible byproduct, but not necessarily something that causes the plan to be safe or correct. This means we need new training strategies that force the connection between thought and action.
Lu: Their future work focuses on two main directions: first, mitigating modality bias through contrastive pre-finetuning to make vision essential, and second, breaking those shortcuts with contrastive learning using negative examples to penalize reliance on 'prior equals plan' shortcuts.
Meng: From an engineering standpoint, making vision indispensable is a big challenge because you have to design the training to enforce that visual input is the only thing that matters for the final decision.
Lalam: I think those contrastive learning ideas are where we can go next, trying to guide the policy away from those easy textual shortcuts and toward genuine causal understanding.
Tom: That's a lot of heavy lifting for future models, but it really highlights that we need causally-aware training methods if we want driving agents that are truly robust.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck