XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
Foundation Model Team, XPeng Inc
XPeng Inc.
cs.AI
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving Summary This paper introduces XCoT-VLA, a Vision-Language-Action (VLA) model for autonomous driving that replaces verbose
Terminology
Summary
XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving
Summary
This paper introduces XCoT-VLA, a Vision-Language-Action (VLA) model for autonomous driving that replaces verbose natural-language Chain-of-Thought (CoT) reasoning with a compact, executable Chain-of-Thought (XCoT) representation. The core motivation is that verbose natural-language CoT is not necessarily an effective interface for control
because it is open-ended, costly to decode, and difficult to optimize as an action-facing representation.
The authors argue that reasoning exposed to the trajectory generator should be decision-critical, retaining only semantics that affect driving behavior; compact, requiring limited autoregressive decoding; and executable, directly conditioning continuous trajectory generation.
Method
XCoT-VLA represents reasoning as a short sequence of 2–6 executable semantic-action tokens (e.g., LEFT TURN PREPARE, DECELERATE, RED LIGHT HOLD) that directly encode decision-critical driving intents and their control-relevant semantics.
These tokens are learned from automatically constructed Reason–Action supervision,
where logged trajectories provide action evidence, while scene context supplies causal semantics such as navigation intent, traffic rules, road structure, and surrounding interactions.
The training-data construction pipeline maps each logged sample (ot, τt) to an executable XCoT sequence zt* through three stages: (i) action-evidence extraction from logged trajectories, (ii) semantic grounding of the action in scene context, and (iii) semantic compression into a canonical XCoT token sequence. The pipeline is formalized as: fground ∘ fact ∘ gXCoT: (ot, τt) → at → (rt*, at) → zt*.
Architecturally, XCoT-VLA uses shared multimodal self-attention
followed by deterministic token-function routing
that applies the Reason FFN to XCoT tokens and the Control FFN to trajectory queries. Specifically, "all valid non-trajectory tokens—including visual, text, task-prompt, ego-status, drive-command, and autoregressive XCoT tokens—are assigned to the Reason FFN. Only the 24 trajectory-query positions are assigned to the Control FFN." This decoupling is formalized with masks satisfying mReason + mControl = 1 and mReason · mControl = 0.
Training occurs in two stages. Stage I uses supervised fine-tuning with a combined objective: LSFT = λXCoT LXCoT + λFM LFM, where LXCoT is a token-level cross-entropy loss on XCoT prediction and LFM is a Conditional Flow Matching loss for trajectory generation. Stage II introduces XCoT Policy Optimization (XCPO), an optional refinement that refines the reasoning policy using trajectory-level rewards while keeping the execution stack fixed
—only the Reason FFN and XCoT prediction head are updated, while the visual encoder, self-attention, Control FFN, and trajectory head remain frozen.
Key Results
The paper reports open-loop planning results on general-distribution and lane-change evaluation sets. On the general set, XCoT-VLA reduces longitudinal ADE-6s from 1.6452 to 1.3233 and FDE-6s-Long from 4.3541 to 3.0887 relative to the trajectory-only SFT baseline. On the lane-change set, lateral ADE-6s decreases from 0.5941 to 0.3091 and FDE-6s-Lat from 1.6160 to 0.6484. The paper states: Verbose CoT does not consistently improve lateral accuracy, while Latent tokens help but remain behind the full Reason–Action supervision of XCoT-VLA.
An ablation study separates the contribution of XCoT supervision from targeted navigation data: XCoT alone reduces ADE-6s-Lat by 18.3% and FDE-6s-Lat by 25.1% relative to the trajectory-only SFT baseline. Adding navigation supervision further reduces ADE-6s-Lat to 0.4518 and FDE-6s-Lat to 1.1122.
For efficiency, XCoT-VLA uses only 2–6 tokens versus 40–80 tokens for verbose CoT. Worst-case reasoning-interface latency on an H100 ranges from 38.6 ms (2K input) to 66.3 ms (6K input) for XCoT-VLA, staying within the 12 Hz (83.3 ms) budget, whereas Verbose CoT exceeds the budget at all input lengths (279.1–306.8 ms).
Training-stability analysis shows that Joint XCoT FT causes a pronounced longitudinal degradation: ADE-6s-Long increases from 1.2997 to 1.6005 and ADE-2s-Long from 0.2774 to 1.1684,
while Decoupled XCoT FT preserves short-horizon longitudinal accuracy and improves ADE-6s-Lat from 0.2302 to 0.1872 and FDE-6s-Lat from 0.5710 to 0.4690.
Qualitative analysis shows XCoT exhibits more proactive and context-aware behaviors, including earlier lane changes, proactive responses to traffic-light constraints, and traffic-adaptive acceleration for efficient maneuvering
compared to the SFT baseline.
Limitations and Future Work
The paper acknowledges that XCoT currently has limited capability for interactive negotiation with surrounding agents, remains sensitive to traffic-rule perception errors, and shows smaller improvements in longitudinal comfort.
Future work will strengthen interaction- and uncertainty-aware reasoning, extend XCoT to richer rule-governed scenarios, and quantitatively evaluate XCPO using closed-loop safety, route-completion, and comfort metrics.
Improvements for AI systems
Based on this paper, I can make the following specific improvements to AI systems:
Improvement: Instead of generating free-form natural language reasoning before action prediction, the AI system uses a compact sequence of 2–6 discrete, decision-critical semantic tokens (e.g., LEFT TURN PREPARE, DECELERATE) that directly condition the action generator. This reduces autoregressive decoding cost by 10–30× and eliminates irrelevant reasoning that doesn't affect control output.
What the improved system can do: Achieve real-time control at 12 Hz on edge hardware (e.g., H100) with 38–66 ms reasoning latency, versus 279–307 ms for verbose CoT—enabling deployment in latency-critical robotics and autonomous driving.
Improvement: Implement deterministic token-function routing where only reasoning tokens (visual, text, XCoT) pass through a Reason FFN, and only trajectory-query positions pass through a separate Control FFN. This prevents joint training interference and stabilizes short-horizon longitudinal accuracy.
Improvement: Automatically construct training labels by extracting action evidence from logged trajectories, grounding it in scene context (navigation, traffic rules, road structure), and compressing into canonical XCoT tokens. This provides dense, causal supervision without manual annotation.
Improvement: After supervised fine-tuning, refine only the reasoning policy (Reason FFN + XCoT head) using trajectory-level rewards, while freezing the control stack. This aligns reasoning with downstream task success without destabilizing action generation.
Improvement: Filter out non-essential reasoning (e.g., descriptive scene narration) and retain only semantics that change driving behavior. Use canonical token compression to reduce redundancy across similar situations.
Improvement: Add XCoT supervision specifically targeting lateral maneuvers (lane changes, turns) that verbose CoT fails to improve. Combine with navigation intent tokens to disambiguate ambiguous road geometry.
Sources
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection