From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "From Representational Complementarity to Dual Systems".
Dev: Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, yet it remains unclear how vision-language models (VLMs) differ internally from standard vision-only encoders,
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: So, this paper "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving" is really digging into the difference between how a vision-language model and a regular vision encoder work internally, especially after they've both been trained to handle driving tasks. I’m curious if this kind of internal difference actually translates to something useful when we take it out of the lab and into messy real-world situations.
Dev: I'm interested in what they claim about that internal structure, Rosa; specifically, how those differences behave once the policy learning process is complete. It seems like a core question for any system relying on these complex backbones to function reliably under stress.
Taro: From an autonomy standpoint, if there are persistent model-specific residuals that survive the diffusion policy, that suggests different models might be suited for distinctly different operational environments when things go wrong in the real world.
Rosa: Exactly, Taro; they’re asking how those VLM and vision-only encoders differ internally and whether those differences survive downstream policy learning. The whole point seems to be figuring out if we can exploit that complementarity.
Dev: And it looks like their investigation focuses on three main questions: representation similarity, behavioral differences in long-tail scenarios, and the resulting system design opportunities. That structure gives us a good roadmap for what they are trying to prove about this paper "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-End Driving".
Taro: I’m particularly interested in the behavioral part, because if the internal representations differ, we need to see if that difference actually manifests when the world gets tricky or unexpected.
Rosa: That's what they find quite interesting; they look at how these representation differences translate into actual driving behavior, which is something we really need to test outside controlled environments.
Dev: The paper mentions analyzing both backbone features and decision-level features using tools like linear CKA and CCA to see where the similarity is happening. That gives us a concrete way to measure those internal relationships mentioned in "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-End Driving".
Taro: So, if the decision level features show more transferability than the backbone features, that’s a big hint about what information is actually useful for the policy when things get complicated.
Rosa: Right, and they also use a Shared–Unique Sparse Autoencoder to check if those shared factors are interchangeable; they found that decision-level features are more transferable across branches than backbone features, but there are still non-transferable residual factors remaining.
Paper summary: Dev: That means the VLM and vision-only encoders aren't just identical after training; there’s a persistent layer of difference that needs to be accounted for in how we use them.
Taro: And those residuals seem to lead directly into the next part of their study, which is looking at how these differences affect actual driving styles in specific situations.
Rosa: Precisely, and that's where they show statistically meaningful behavioral differences, such as vision-only policies being more conservative while VLM policies are more assertive in aggregate.
Dev: But the most compelling finding seems to be this long-tail phenomenon where the complementarity shows up decisively in specific subsets of scenarios.
Taro: That idea of a "long-tail phenomenon" suggests that the advantages might not be general but tied to very specific, complex interactions or dense clutter cases where semantic understanding really helps.
Rosa: It’s a critical point because if we can identify those tricky scenarios, we could potentially design systems that switch between the branches based on what’s happening visually.
Dev: And that leads into their final section where they present two specific system designs, HybridDriveVLA and DualDriveVLA, designed to exploit this complementarity rather than just comparing the models statically.
Taro: Those are the practical applications; seeing how they build a hybrid system or a fast-slow variant based on these findings shows how theory can translate into something we could actually deploy in a vehicle.
Rosa: Yes, and those systems show measurable performance gains, like HybridDriveVLA improving PDMS from ninety point eight zero to ninety-two point one zero without changing the policy training itself, which is very neat for deployment considerations <ref:2602.10719#pg2>.
Dev: And DualDriveVLA addresses latency by using the vision-only branch as a default and only calling in the VLM when necessary, achieving about a one point nine times lower-latency speedup over the VLM baseline for a PDMS of ninety-one point zero zero.
Taro: So, what they’re showing us is that we can use this analysis to create smarter decision-making logic for systems that need to operate reliably across a huge variety of driving conditions, not just the average case.
Rosa: It really feels like the whole idea of "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving" is about finding a way to leverage the strengths of both types of encoders intelligently.
Dev: And it’s important that they point out their limitation, which is that even with these analyses, when they test rule-based or learned gates built from alignment statistics or latent features, the gains remain marginal, with the best PDMS only reaching ninety point eight zero to nine <ref:2602.10719#pg2,learned gates built from alignment statistics or latent features, the gains remain>.
Taro: That means it’s not just about having a VLM and a vision-only encoder; it’s about figuring out the right way to connect them structurally rather than just hoping they work together automatically.
Paper summary: Rosa: It suggests that for real-world application, we need to be careful; representation-only gating isn't enough on its own to reliably predict the trajectory quality in complex driving situations.
Dev: So, while the analysis is strong on identifying where the differences lie, it points toward a need for more sophisticated decision-making mechanisms than just simple feature comparison to actually get those gains consistently.
Taro: Looking ahead, this work opens up avenues for designing more adaptive autonomy systems that can dynamically switch between different processing modes based on real-time scene complexity detected by the encoders.
Rosa: And it makes me wonder how long these dual system approaches will be viable once we move from simulated driving to actual road testing; it’s a big question for field robotics.
Dev: The latency improvements shown with DualDriveVLA are promising, but we need to ensure that the decision-making overhead introduced by the switching mechanism doesn't create new failure modes in our real-time loop.
Taro: If we can reliably predict which branch to use based on scene cues, then adapting that logic for unpredictable human behavior on the road seems like a viable path forward for robust autonomy.
Rosa: So, in short, this paper provides a deep look into why VLMs and vision-only encoders aren't just redundant after policy learning, showing that their differences can be leveraged through specific architectural choices.
Dev: It’s an analysis-driven account of why VLM and vision-only policies are not redundant after policy learning, which is what the paper "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-End Driving" aims to provide.
Taro: The implication here is that we shouldn't treat backbones as black boxes; we need tools to understand their internal representation geometry before we try to use them in high-stakes autonomy.
Rosa: That’s a big shift in how I think about integrating these models into my field projects, moving from just plugging in a model to understanding its specific strengths and weaknesses under different conditions.
Dev: And for the control side, the idea of having both branches available and using a learned scorer to pick the best path sounds like a solid engineering starting point for improving our planning metrics.
Taro: The long-term impact could be in creating autonomy that is inherently more robust because it can adapt its internal reasoning based on the perceived difficulty of the environment.
Rosa: It certainly gives us a lot to think about regarding how we design these systems to handle the unpredictable nature of driving outside of a controlled simulation setting.
Conclusion: Rosa: So, we've looked at how this paper shows that VLM and vision-only policies aren't just redundant after they’ve been trained to drive, and now we’re getting to the conclusion on "From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving."
Dev: I agree, Rosa; the title really captures the essence of what they did, showing how those two different kinds of encoders can actually work together. The authors are doing some heavy lifting here by analyzing their internal representations and then designing systems to use that difference.
Taro: I think it’s important to focus on what this means for autonomy; if we can harness these specific strengths, we might be able to build systems that handle a much wider range of driving situations than current models allow.
Rosa: That’s what I’m thinking; the core idea is that there are different kinds of driving scenarios where one model excels and the other is better, and they found a way to combine them for better overall performance.
Dev: From an engineering standpoint, this suggests we can design more adaptive control loops that switch between processing modes depending on what the visual input looks like right now.
Taro: I'm really interested in how this translates to real-world reliability; if the system can intelligently choose the right tool for a tricky situation, that’s huge for handling unexpected events on the road.
Rosa: That brings up a big question for me; can we actually rely on these hybrid systems to perform reliably outside of a perfectly controlled lab environment, and how long do you think that reliability lasts?
Dev: The latency improvements they showed with their DualDriveVLA variant are promising, but we need to be sure that the decision-making overhead doesn't introduce new failure modes in a fast control loop.
Taro: If we can reliably predict which branch to use based on scene complexity, then adapting that logic for unpredictable human behavior on the road seems like a viable path forward for robust autonomy.
Rosa: It certainly gives us a lot to think about regarding how we design these systems to handle the unpredictable nature of driving outside of a controlled simulation setting.
Dev: And remember, while their analysis is strong on identifying where the differences lie, they did flag that representation-only gating doesn't always reliably predict trajectory quality when things get really complex.
Taro: That limitation is important because it tells us we need to go beyond just comparing feature alignments; we need a better way to use that knowledge to select the right model.
Rosa: So, the main point here is that this research offers an analysis-driven account of why these models are different and provides concrete system designs for leveraging those differences.
Dev: Exactly, it’s not just about having two encoders; it’s about structuring how they interact to get a measurable performance boost in specific challenging cases.
Taro: This work opens up avenues for designing more adaptive autonomy systems that can dynamically switch between different processing modes based on the perceived difficulty of the environment.
Rosa: It makes me wonder how long these dual system approaches will be viable once we move from simulated driving to actual road testing; it’s a big question for field robotics.
Institute for AI Industry Research (AIR), Tsinghua University
cs.RO, cs.CV
Submitted: 2026-02-11
Updated: 2026-10-02
Comments: Accepted at NeurIPS 2026
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 90/100
The gist: Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, yet it remains unclear how vision-language models (VLMs) differ internally from
Key concepts
- Representation Geometry
- This refers to how the internal mathematical structure or 'shape' of features extracted by different vision models relates to each other. The researchers measured this using techniques like CKA and CCA to see if VLM and vision-only encoders share similar underlying feature spaces during policy learning.
- Decision-Level Features
- These are the features derived from individual decisions made by the policy, rather than the raw input images. The analysis found these features are more transferable between VLM and vision-only branches than backbone features, suggesting that high-level behavioral insights are more compatible across different model architectures.
- Long-Tail Phenomenon
- This describes a situation where success or performance differences occur only in specific, rare subsets of scenarios rather than uniformly across all cases. The paper found that VLM policies are significantly stronger in these 'long-tail' cases, such as complex interactions or dense clutter, demonstrating where their unique capabilities shine.
- HybridDriveVLA
- This is a system design approach that runs both the vision-only and VLM branches simultaneously. It uses a learned trajectory scorer to intelligently select the best path by interpolating between the two models, resulting in improved performance without needing to retrain the main policy.
Terminology
Summary
Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, yet it remains unclear how vision-language models (VLMs) differ internally from standard vision-only encoders, and whether such differences survive downstream policy learning. This work investigates the relationship between VLM and vision-only policies by analyzing their representation geometry, behavioral differences in long-tail scenarios, and the resulting system design opportunities to exploit this complementarity.
How it works
The study is organized around three progressive research questions: RQ1 (Representation), RQ2 (Behavior), and RQ3 (System). The analysis is conducted under a unified VLM-hidden + diffusion-policy paradigm, comparing multiple VLM families against commonly used vision-only encoders like ResNet, ViT, and EVA-CLIP.
RQ1 focuses on representation similarity. Using linear CKA and CCA, the researchers found that policy learning substantially enlarges the shared subspace between the two branches
but does not collapse them into a single redundant representation. They utilized a Shared–Unique Sparse Autoencoder (SAE) to determine if shared factors are functionally interchangeable; they found that decision-level features are more transferable across branches than backbone features, yet still not fully interchangeable,
indicating that non-transferable residual factors remain.
RQ2 investigates behavioral differences. The analysis shows that these residual representational differences translate into statistically meaningful behavior. At the aggregate level, vision-only policies are relatively more conservative, whereas VLM policies are more assertive.
Crucially, the paper finds complementarity is a long-tail phenomenon,
where decisive wins occur in specific subsets of scenarios; for instance, VLM dominates in semantically complex or interaction-heavy scenarios
and occlusion/dense clutter cases.
RQ3 translates this complementarity into practical gains. The researchers introduce two lightweight systems to exploit the findings:
-
HybridDriveVLA: This runs both branches and uses a
learned trajectory scorer for selection,
improving PDMS from 90.80 to 92.10 without changing policy training. -
DualDriveVLA: This is a fast–slow variant that runs the vision-only branch by default and invokes the VLM only when necessary, achieving 91.00 PDMS with about
1.9× lower-latency speedup over the VLM baseline.
Analysis of Representation and Behavior
RQ1 analysis utilized paired features extracted from driving scenarios at two levels: (i) backbone features, where CKA rose from about 0.22 to 0.54 after the shared diffusion planner, and (ii) decision-level features, which showed even greater alignment (CKA rising from about 0.22 to 0.54). The Shared–Unique SAE confirmed that while policy learning reduces the self–cross gap
(∆cross), indicating more transferable factors, non-transferable residual factors remain,
explaining why replacing a VLM encoder with a vision-only one causes only a moderate average PDMS drop.
RQ2 quantified behavioral differences by defining per-scenario advantage metrics. Under conservative counting, the two branches show distinct driving styles: vision-only encoders are relatively more effective in simple, geometry-dominant scenes,
while VLMs are substantially stronger in long-tail, semantically complex, and interaction-heavy cases.
This is evidenced by a decisive win subset where VLM wins 176 scenarios versus 103 for ViT when using a strict threshold of τ = 0.9. Furthermore, trajectory statistics show the VLM branch tends to advance further while also reacting with stronger braking-side corrections when required.
System Design and Exploitation
RQ3 focused on converting complementarity into measurable gains. HybridDriveVLA constructs an 11-trajectory candidate set
by interpolating between the two branches along a cross-model style axis and selecting the final trajectory using a learned trajectory scorer trained to predict PDMS sub-score components. DualDriveVLA operationalizes this by using a threshold γ: it runs the vision-only branch first; if its predicted meta-score is insufficient, it invokes the VLM branch, achieving a 1.9× latency speedup while maintaining high accuracy (91.00 PDMS).
Conclusion and Takeaways
The overall contribution is an analysis-driven account of why VLM and vision-only policies are not redundant after policy learning.
The key takeaways include:
-
Policy learning enlarges the shared subspace but leaves non-transferable residuals.
-
Decision-level features are more transferable than backbone features, though not fully interchangeable.
-
Representation-only gating fails to reliably predict trajectory quality, motivating behavioral analysis instead.
Improvements for AI systems
As a fastidious researcher, I have analyzed this paper, From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving,
focusing on its core findings regarding the complementarity between Vision-Language Models (VLMs) and vision-only encoders in end-to-end planning.
The paper suggests that the primary improvement lies not in simply using a VLM, but in intelligently exploiting the residual differences between a VLM and a vision-only backbone.
Here are the specific improvements for AI systems based on this research:
)
)
)
)
-
Exploiting Complementarity via Hybrid/Dual Architectures (RQ3): Implement a
HybridDriveVLA
system that runs both the VLM branch and a vision-only branch simultaneously, interpolates between their candidate trajectories along the cross-model style axis, and uses a learned trajectory scorer to select the optimal path. -
Adaptive Fast/Slow Deployment (DualDriveVLA): Develop a
DualDriveVLA
system where the default path is generated by the faster vision-only branch. The VLM branch is invoked only when the fast-path score falls below a learned threshold, achieving significant latency reduction (up to 1.9× speedup) while maintaining high accuracy (PDMS of 91.00). -
Behavioral Strategy Selection Based on Scenario Complexity (RQ2): Instead of static representation gating, implement a decision mechanism that dynamically selects the backbone based on scenario-level characteristics (e.g., using semantic grouping or complexity indicators derived from SAE energy decomposition) to choose between the VLM and vision-only branch.
-
Trajectory-Level Scoring for Selection: The system should shift selection criteria from static representation cues to dynamic, trajectory-level signals. This involves training a lightweight scorer (e.g., using a DINOv2 feature extractor) to predict PDMS sub-scores (safety, progress, comfort) and selecting the candidate trajectory that maximizes this meta-score.
-
Cross-Model Candidate Construction: The system should construct a small, interpretable candidate set by interpolating between the VLM and vision-only branches along a learned style axis (e.g., using interpolation factors from 0.1 to 0.9) rather than relying solely on sampling from a single model.
The improved AI system can achieve the following capabilities:
-
Enhanced Robustness in Complex Scenarios: The system will exhibit superior performance (higher PDMS) in
long-tail
scenarios, such as intersections, complex merges, and dense clutter/occlusion, where the VLM's semantic reasoning provides a decisive advantage over geometric-dominant vision-only features. -
Optimized Accuracy with Reduced Latency: By employing DualDriveVLA, the system can achieve near-VLM accuracy (91.00 PDMS) while operating at significantly lower inference latency (150 ms), making it suitable for real-time, high-throughput autonomous driving applications.
-
Style-Aware Driving: The system will exhibit distinct driving styles—being more
assertive
in semantically complex situations and moreconservative
in simple geometry scenarios—allowing the planner to adapt its risk profile based on the perceived scenario difficulty. -
Model-Agnostic Decision Making: By focusing on trajectory-level scores rather than backbone features, the system becomes less brittle; it can reliably select the best path even when static representation cues (like simple CKA similarity) are misleading due to feature distribution differences.
-
Efficient Resource Utilization: The fast/slow deployment strategy allows for efficient resource management, only engaging the computationally expensive VLM branch when necessary, leading to substantial computational savings compared to running a full VLM-based policy in every instance.
Abstract
Vision-Language-Action (VLA) driving augments end-to-end (E2E) planning with language-enabled visual backbones, but how VLM representations differ from vision-only encoders after policy learning, and whether such differences matter for planning, remains unclear. Under a unified VLM-hidden + diffusion-policy paradigm, we compare multiple VLM families/scales (InternVL3 and Qwen3VL) with standard vision-only encoders (ResNet, ViT, and EVA-CLIP) while keeping the downstream planner fixed. We study representation, behavior, and system design. CKA/CCA and Shared--Unique SAE show that policy learning enlarges a common decision subspace, but both branches retain non-transferable residual factors. We further replicate this shared-plus-unique representation pattern on nuPlan using the AsyncDrive planning stack. Latent-intervention policies and scenario-level analysis further show that these residuals are behaviorally meaningful: vision-only encoders are stronger in simple geometry-dominant scenes, whereas VLMs are more effective in semantically complex and interaction-heavy long-tail cases. The two branches also exhibit distinct progress--braking and path-choice tendencies, and an oracle best-of-two VLM+ViT selector reaches 93.58 PDMS on NAVSIM. We convert this complementarity into two lightweight systems: HybridDriveVLA, which selects from a compact cross-model candidate set using a learned trajectory scorer and improves PDMS from 90.80 to 92.10, and DualDriveVLA, a fast--slow variant that invokes the VLM in only 15% of scenarios, achieving 91.00 PDMS with about 1.9 times lower latency than the VLM baseline. Code will be available at https://github.com/WilliamXuanYu/HybridDriveVLA.
Sources
- Is a 3D-Tokenized LLM the Key to Reliable Autonomous Driving?
- Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
- DrivingGPT: Unifying Driving World Modeling and Planning with Multi-modal Autoregressive Transformers
- DriveLM: Driving with Graph Visual Question Answering
- ARTEMIS: Autoregressive End-to-End Trajectory Planning with Mixture of Experts for Autonomous Driving
- ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
- iPad: Iterative Proposal-centric End-to-End Autonomous Driving
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
- Driving on Registers
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- End-to-End Driving with Online Trajectory Evaluation via BEV World Model
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
- DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
- GPT-Driver: Learning to Drive with GPT
- A Language Agent for Autonomous Driving
- Reason2Drive: Towards Interpretable and Chain-based Reasoning for Autonomous Driving
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving