BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Bing Zhan, Shuyao Shang, Shuo Lu, Yuan Xu, Zhao Wang, Yida Wang, Xueyang Zhang, Kun Zhan, Jiahao Gu
NLPR, Institute of Automation, Chinese Academy of Sciences · Li Auto Inc.
cs.RO, cs.AI, cs.CV
Submitted: 2026-08-19
Updated: 2026-08-20
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving Abstract Autonomous driving requires planning under both semantic constraints and predictive
Terminology
Summary
BrainWAM: Action-Space Coordination of Semantic Priors and Predictive Dynamics for Autonomous Driving
Abstract
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling. This naturally motivates a unified planner that can leverage both semantic priors and predictive dynamics. However, we find that a naive combination through joint token-level attention suffers from an attention-allocation mismatch, where semantic shortcuts dominate the shared attention space and suppress predictive dynamics. Inspired by neuroscience evidence that complex behavior arises from coordination among functionally specialized systems, we propose BrainWAM, a structured action-space coordination framework that converts semantic reasoning and predictive world modeling into two specialized action-oriented pathways, and aligns them at the level of compact action representations. We further introduce an asynchronous rectified-flow inference strategy with decoupled video and action denoising, which shortens inference latency while preserving planning-relevant predictive context. BrainWAM reaches state-of-the-art performance on both NAVSIM v1 (89.5 PDMS) and NAVSIM v2 (89.6 EPDMS), consistently outperforming VLA-only or WAM-only methods, highlighting BrainWAM as a practical and promising direction for autonomous driving systems.
Introduction
Autonomous driving requires planning under two tightly coupled forms of evidence: semantic constraints and predictive dynamics. VLA models leverage the world-knowledge priors of VLMs, making them effective at grounding observations in traffic rules, route instructions, scene semantics, and high-level driving intent. WAMs, inspired by recent progress in action-conditioned world modeling, instead learn how actions and future states evolve together, providing predictive context for motion trends, interaction outcomes, and physical feasibility. These strengths are naturally complementary: VLA models provide task-aware semantic and decision priors but usually lack explicit modeling of future scene evolution, whereas WAMs provide future-aware dynamics and physical priors but are less reliable at rule-aware and intent-driven reasoning.
A common direct design is Tri-modal Joint Attention (Tri-MoT), which places VLM tokens, Video Generative Model (VGM) tokens, and action tokens into one shared attention space. However, we find that this raw-token fusion can even underperform WAM alone. To diagnose this issue, we visualize how action tokens attend to VLM and VGM tokens. As shown in Fig. 2, action tokens attend more strongly to semantic-level VLM tokens than to pixel-level VGM tokens across most Transformer layers, especially in shallow layers. This asymmetry follows the modality competition observed in joint multimodal training, where the modality that is easier to learn dominates optimization and suppresses the other. Here the clean and semantically compact VLM tokens are easier to learn, while the VGM tokens are still being denoised and provide lower-signal features, so action tokens take the VLM shortcut and underuse the predictive video tokens. As a result, directly mixing high-dimensional heterogeneous tokens induces an attention-allocation mismatch: semantic signals dominate the shared interaction space and weaken the predictive dynamics needed for planning.
To address this challenge, we draw inspiration from neuroscience: complex behavior emerges not from homogenizing all signals into one undifferentiated representation, but from coordination among functionally specialized systems. The left hemisphere is often associated with language, symbolic, and sequential processing, while the right hemisphere plays an important role in visuospatial and holistic understanding; the two hemispheres exchange information through the corpus callosum, and motor intent is further coordinated and refined by the cerebellum. This organization suggests a computational principle for VLA-WAM integration: semantic reasoning and predictive world modeling should first develop specialized, behavior-relevant action representations, and then coordinate through compact action-level communication.
Motivated by this principle, we propose BrainWAM, a brain-inspired action-space coordination framework for autonomous driving. BrainWAM structures semantic reasoning and predictive world modeling into two complementary action-oriented pathways: a left-hemisphere pathway distills traffic-scene semantics, route instructions, and rule-aware decision priors from VLM, while a right-hemisphere pathway distills spatiotemporal dynamics, physical consistency, and future-interaction cues from VGM. The two pathways communicate bidirectionally over compact action tokens through a corpus-callosum-inspired Callosal Action Bridge (CAB), and a cerebellum-inspired Cerebellar Intent Fusion (CIF) module coordinates the refined action intents and decodes them into an executable trajectory.
Contributions
-
We propose BrainWAM, an action-level coordination framework that combines VLM-based semantic reasoning with WAM-based predictive world modeling. Inspired by brain functional specialization, BrainWAM converts instruction-aware semantic constraints and future-dynamics priors into complementary action representations, and coordinates them in a unified action space.
-
We identify an attention-allocation mismatch in Tri-MoT: its action tokens attend disproportionately to semantic tokens in most layers, causing raw-token fusion to underperform WAM-only planning.
-
BrainWAM achieves 89.5 PDMS on NAVSIM v1 and 89.6 EPDMS on NAVSIM v2, outperforming strong end-to-end driving, VLA-based, and world-model-based methods.
Method
The method is trained in three stages. First, the WAM branch learns prediction-grounded action representations from future scene dynamics. Second, the VLA branch learns semantic-grounded action representations from visual observations and language instructions. Finally, both branches are frozen, while CAB, CIF, and the final action decoder are trained to coordinate the two action streams and generate the final trajectory.
WAM Branch. The WAM branch learns prediction-grounded action representations from future scene prediction. Given the current observation, a video generative backbone predicts future visual latents, while a rectified-flow action expert generates the ego trajectory. We perturb the video latent and the action trajectory with independent rectified-flow timesteps. This decoupled schedule allows the video stream to terminate early after forming predictive context, while the action stream continues denoising to generate the trajectory. We adopt Wan2.2-TI2V-5B as the video backbone and attach a lightweight action expert. The video backbone performs denoising over future video latents and produces visual tokens that capture scene dynamics, while the action expert performs trajectory denoising and produces action tokens. Dual-MoT modules couple the two streams through shared self-attention, enabling visual dynamics and action trajectories to interact, while modality-specific feed-forward networks preserve their distinct modeling capacities. The branch predicts video and action vector fields following Flow Matching and Rectified Flow, which define a linear path between clean data and Gaussian noise.
VLA Branch. The VLA branch uses a VLM backbone to extract semantics from visual observations and language instructions. It complements the WAM branch through scene-level understanding and route-conditioned intent rather than future visual prediction. A rectified-flow action expert converts VLM features into semantic-grounded action representations. We adopt Qwen3-VL-4B as the VLM backbone and equip it with a lightweight action expert for trajectory modeling. The VLM encodes multi-view images and driving instructions into semantic tokens, and ego history into state tokens. The action expert processes the noisy trajectory into action tokens. Dual-MoT modules couple semantic, state, and action tokens through shared self-attention to guide action denoising.
Joint Training with CAB and CIF. The joint stage couples the pretrained WAM and VLA branches. Both branches are frozen, while only CAB, CIF, and the final action decoder are optimized. This preserves their pretrained modeling capabilities and focuses learning on cross-stream coordination. Both action experts receive the same noisy trajectory and action timestep, placing the action tokens from both streams at the same noise level, while the WAM video branch retains its own timestep to provide predictive context.
Callosal Action Bridge (CAB): Inspired by the corpus callosum connecting specialized hemispheres, CAB enables bidirectional interaction between prediction-grounded action tokens and semantic-grounded action tokens. Unlike token-level fusion, CAB avoids mixing raw VLM and video tokens within a shared attention pool. At each layer, CAB computes bidirectional cross-stream messages using cross-attention, and the two action streams are subsequently updated through gated residual injection with learnable residual gates that are zero-initialized, preserving pretrained action streams initially while learning cross-stream updates during joint training.
Cerebellar Intent Fusion (CIF): Inspired by the cerebellum's role in motor coordination, CIF integrates the refined action streams into a unified representation. It concatenates both streams, processes them with a lightweight Transformer module, and averages the resulting outputs. The fused representation is decoded into the action velocity. Joint training supervises only the fused prediction.
Experiments
Benchmark and Datasets. We evaluate planning performance on NAVSIM v1 and NAVSIM v2. NAVSIM is built upon OpenScene, a reprocessed version of nuPlan, and consists of real-world driving logs. At each frame, the model predicts a 4-second trajectory at 2Hz, yielding 8 waypoints. The predicted trajectory is evaluated in a short-horizon, non-reactive simulation. NAVSIM v1 reports the Predictive Driver Model Score (PDMS), which aggregates No at-fault Collision (NC), Drivable Area Compliance (DAC), Time-To-Collision (TTC), Comfort (C), and Ego Progress (EP). NAVSIM v2 extends this metric with two additional penalty multipliers, Driving Direction Compliance (DDC) and Traffic Light Compliance (TLC), and three weighted subscores, Lane Keeping (LK), History Comfort (HC), and Extended Comfort (EC), reporting the Extended PDMS (EPDMS).
Implementation Details. Each of the three stages is trained for 100K steps on 8 NVIDIA H20 GPUs with a per-GPU batch size of 6. We use AdamW with a cosine learning-rate schedule, 200 warmup steps, and a peak learning rate of 5 × 10−5. Training uses bf16 mixed precision, with checkpoints saved every 3K steps. At inference, we use 3-step rectified-flow sampling for the action streams.
Main Results. On NAVSIM v1, BrainWAM achieves a PDMS of 89.5, outperforming both VLA-based and world-model-based baselines. The gains are most pronounced in DAC and EP, indicating improved drivable-area compliance and driving progress, while maintaining competitive NC, TTC, and comfort scores. On NAVSIM v2, BrainWAM achieves state-of-the-art performance with an EPDMS of 89.6. The improvements are primarily driven by EP and EC, whereas several rule-compliance metrics are already near saturation.
Further Analysis and Ablation Studies.
-
Branch complementarity: WAM-only achieves 88.1 PDMS and substantially outperforms VLA-only (86.1 PDMS), demonstrating the strong planning prior provided by predictive modeling on NAVSIM. BrainWAM further improves PDMS to 89.5, exceeding both single-branch variants, suggesting that semantic and predictive action representations provide complementary information under action-level coordination.
-
Action-level coordination vs. token-level fusion: Tri-MoT achieves 87.8 PDMS, underperforming the WAM-only variant. This indicates that directly mixing VLM and video tokens in a shared attention space does not effectively transfer semantic knowledge to planning. By keeping raw modality tokens separate and interacting only through action representations, BrainWAM improves PDMS to 89.5. Because both methods use identical backbones and comparable parameter counts, the gain is attributable to the coordination mechanism rather than increased model capacity.
-
Effectiveness of CAB and CIF: Using CAB or CIF alone yields 88.7 and 88.5 PDMS, respectively, whereas combining them increases PDMS to 89.5. The improvement is concentrated in DAC and EP, while NC and TTC remain stable. These results suggest that CAB facilitates intermediate interaction between the two action streams, while CIF consolidates their final representations.
-
Asynchronous video denoising: With no video denoising, the model loses predictive context and drops to 79.3 PDMS and 75.8 EPDMS, confirming that video dynamics are essential to planning. A single video step restores PDMS to 89.3, after which performance remains between 89.3 and 89.5 as the number of steps increases to 3, while latency rises from 475 ms to 644 ms. Thus, one early video step provides most of the useful predictive context, offering a favorable trade-off between accuracy and efficiency.
Qualitative Analysis. The paper compares VLA-only, WAM-only, and BrainWAM across representative scenarios including navigation following, red-light understanding, interactive negotiation, and trajectory feasibility on curved roads. VLA-only handles semantic-grounding challenges better than WAM-only, while WAM-only performs better in cases requiring anticipation of scene evolution and physical consistency. BrainWAM handles all four cases by coordinating semantic-grounded and prediction-grounded action representations, reducing the failure modes observed in the single-branch variants.
Conclusion
In this work, we study how to effectively combine VLA-based semantic reasoning and WAM-based predictive world modeling for end-to-end autonomous driving. We first reveal that naive shared-token fusion suffers from an attention-allocation mismatch: action tokens attend disproportionately to semantic tokens, which weakens the predictive signals provided by the world model, resulting in suboptimal planning. Motivated by this, we propose BrainWAM, which allows the two branches to first produce semantic-grounded and prediction-grounded action representations, and then coordinate them through structured interaction in a unified action space, while preserving the complementary specialization of semantic and predictive pathways. BrainWAM achieves state-of-the-art performance on both NAVSIM v1 and v2, demonstrating its effectiveness and potential for autonomous driving systems.
Appendix Highlights
-
Modality Imbalance in Tri-MoT: The attention imbalance is related to modality competition in multimodal learning. When heterogeneous modalities are jointly optimized in a shared representation space, the model tends to rely on the modality that offers more stable and easily optimizable signals. The VLM tokens are clean and stable, while VGM tokens are produced by a rectified-flow denoising process and are less stable. Action tokens take the VLM tokens as the easier-to-learn modality and assign them higher attention. Two observations support this: (1) when video denoising is disabled, PDMS drops to 79.3, confirming predictive context is essential; (2) adding VLM tokens to the shared attention space does not help, as Tri-MoT reaches only 87.8 PDMS, below WAM-only (88.1), even though it has access to strictly more information.
-
CAB Implementation: CAB operates on two action-token streams, each containing 8 tokens with a hidden dimension of 1024. CAB is inserted at Layers 9 and 18 of the two action experts. Each CAB contains two parallel multi-head cross-attention modules with 8 heads and a head dimension of 128. The cross-attention output is injected through a gated residual with zero-initialized gates. Two CAB blocks contain approximately 16.8M parameters. Ablations show that two CAB blocks are sufficient, as increasing to 3, 5, or 28 yields comparable performance (89.2–89.3 PDMS).
-
CIF Implementation: CIF operates on two action streams, each containing 8 tokens with a hidden dimension of 1024. The two streams are projected separately to a shared 1024-dimensional space, with a learnable source embedding. The concatenated sequence is processed by a 2-layer Transformer with 8 attention heads using action-timestep-conditioned AdaLN modulation. CIF contains approximately 49.3M parameters. Ablations show that Transformer-based fusion performs best (89.3 PDMS) compared to MLP (88.8) and gated fusion (89.1), and two Transformer layers are sufficient.
-
Freezing Pretrained Branches: Full-model fine-tuning obtains 88.8 PDMS, whereas selectively updating CAB, CIF, and the action decoder improves PDMS to 89.5. The WAM and VLA branches exhibit different convergence speeds during independent training (VLA-only reaches 86.1 PDMS after 54K steps, while WAM-only requires 81K steps to reach 88.1 PDMS). Freezing the two pretrained branches avoids optimization imbalance and provides stable inputs to CAB and CIF.
-
Limitations: BrainWAM jointly executes the WAM and VLA branches and retains a generative video backbone during inference. Consequently, its computational and memory costs remain higher than those of a single-branch planner. Although the asynchronous denoising schedule reduces inference latency to 475–644 ms, this runtime does not yet satisfy the strict real-time requirements of practical in-vehicle deployment. Further efficiency improvements may require compressing or distilling the video branch, reducing redundant computation between the two pathways, and developing more aggressive feature-reuse or early-exit strategies.
Improvements for AI systems
Improvements to AI Systems Based on BrainWAM:
- Action-Space Coordination Module for Multi-Modal Fusion
-
Instead of naive token-level attention fusion (which suffers from modality competition), implement a two-pathway architecture where each modality (semantic/textual and predictive/visual-dynamics) first produces its own compact, task-relevant action representations independently.
-
Add a gated cross-attention bridge (CAB) between these action streams, with zero-initialized residual gates to preserve pretrained capabilities initially and gradually learn cross-stream coordination.
-
Add a final fusion transformer (CIF) that consolidates both streams into a single executable output.
-
Resulting capability: AI systems can combine heterogeneous modalities (e.g., language + video, or rules + physics) without one modality dominating or suppressing the other, leading to more balanced and robust decision-making.
- Asynchronous Denoising Schedule for Generative Models
-
Decouple the denoising timesteps of different generative streams (e.g., video prediction vs. action generation) during inference.
-
Allow the predictive context stream (e.g., video) to terminate early after forming sufficient context, while the primary output stream (e.g., action) continues denoising.
-
Use a single early step for the auxiliary stream to capture most of the useful predictive information.
-
Resulting capability: Reduces inference latency by 20–30% (e.g., from 644ms to 475ms) while retaining 99.8% of the performance, enabling near-real-time deployment of generative world models in latency-sensitive applications like robotics or autonomous driving.
- Freezing Pretrained Branches During Joint Coordination Training
-
When combining two specialized pretrained models (e.g., a VLM and a world model), freeze both backbones and only train the coordination modules (bridge, fusion, and output decoder).
-
This avoids optimization imbalance caused by differing convergence speeds between branches.
-
Resulting capability: Enables stable and efficient integration of large pretrained models without catastrophic forgetting or performance degradation, improving final accuracy by 0.7–1.0% over full fine-tuning.
- Modality-Competition-Aware Attention Regularization
-
During joint training of heterogeneous modalities, monitor attention allocation across layers.
-
If one modality consistently receives disproportionately high attention (e.g., semantic tokens over predictive tokens), restructure the architecture to keep raw modality tokens separate and only allow interaction at the level of compact task-relevant representations.
-
Resulting capability: Prevents the
shortcut
problem where easier-to-learn modalities suppress harder but valuable signals, improving performance by 1.7 PDMS over naive fusion (89.5 vs. 87.8) while using identical backbones.
- Dual-Pathway Specialization with Complementary Supervision
-
Train each pathway (semantic and predictive) independently with its own objective before joint coordination: one learns from language-grounded semantics, the other from future-state prediction.
-
Then freeze both and train only the coordination layers.
-
Resulting capability: Each pathway develops deep, specialized expertise without interference, and the final system benefits from both rule-aware reasoning and physical-feasibility prediction, outperforming either single pathway by 1.4–3.4 PDMS.
- Cross-Stream Action Token Alignment at Equal Noise Levels
-
During joint coordination, feed the same noisy trajectory and same action timestep to both action experts, ensuring their action tokens are at identical noise levels.
-
This creates a fair and comparable action space for cross-stream communication.
-
Resulting capability: Enables meaningful bidirectional interaction between semantic and predictive action representations, improving coordination quality and final output accuracy.
- Scalable Coordination with Minimal Parameter Overhead
-
Use only 2–3 bridge blocks (16.8M params) and a 2-layer fusion transformer (49.3M params) for coordination, rather than full joint attention over all tokens.
-
Resulting capability: Achieves state-of-the-art performance with only 66M additional parameters, making the approach feasible for systems with limited compute or memory budgets.
Sources
- Qwen3-VL Technical Report
- NuPlan: A closed-loop ML-based planning benchmark for autonomous vehicles
- Pseudo-Simulation for Autonomous Driving
- WorldVLA: Towards Autoregressive Action World Model
- GAIA-1: A Generative World Model for Autonomous Driving
- Map-World: Masked Action planning and Path-Integral World Model for Autonomous Driving
- ADriver-I: A General World Model for Autonomous Driving
- Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- DynVLA: Learning World Dynamics for Action Reasoning in Autonomous Driving
- DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Unified Vision-Language-Action Model
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning
- Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving