FlashDrive: Flash Vision-Language-Action Inference for Autonomous Driving

arXiv:2608.12932 · cs.AI · Submitted 2026-08-13 · Read on arXiv

Zekai Li, Yihao Liang, Hongfei Zhang, Jian Chen, Yesheng Liang, Zhijian Liu

University of California San Diego · Princeton University · Independent Researcher

cs.AI

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 15 pages; 8 figures

Code: https://github.com/z-lab/flashdrive

Project page: https://z-lab.ai/projects/flashdrive

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 95/100

The gist: FlashDrive is an algorithm-system co-design framework that targets all four stages of Vision-Language-Action (VLA) model inference simultaneously to bring end-to-end autonomous driving substantially

Terminology

Summary

FlashDrive is an algorithm-system co-design framework that targets all four stages of Vision-Language-Action (VLA) model inference simultaneously to bring end-to-end autonomous driving substantially closer to real-time deployment. The paper identifies that VLA inference is not a single bottleneck but a cascade of four distinct stages: visual encoding, language-model prefill, autoregressive reasoning token generation, and flow-matching denoising for action prediction.

The core observation is that each stage harbors a distinct form of redundancy, and each admits a correspondingly distinct algorithmic shortcut. The four stage-specific techniques are:

Streaming Inference (Encode & Prefill): In continuous driving, the VLA model processes a sliding window of temporal frames (typically 4 frames × 4 views) that advances by one frame per timestep, so three out of four frames were already encoded at the previous step. Re-encoding them from scratch wastes ∼75% of the visual computation. FlashDrive encodes only the newest frame and persists the KV cache from preceding frames. To handle view-major token ordering, incoming frames are inserted at the end of each camera view with a streaming attention mask that preserves cross-view causality. Since visual token positions shift with each new frame, keys are cached pre-RoPE and rotary embeddings applied on the fly. This reduces the effective sequence length by 75%, yielding over 3× speedups in both encode and prefill. A lightweight streaming fine-tuning procedure adapts the action expert to the resulting distributional shift, recovering accuracy from 2.04 m minADE1 (without fine-tuning) to 1.73 m minADE1.

Speculative Reasoning (Decode): Autoregressive reasoning is the largest single latency contributor, accounting for 271.7 ms (37.9% of total latency) at just 56.4 tokens-per-second throughput. The paper observes that driving-domain reasoning chains are short (16 tokens), follow a highly structured template, and have substantially lower per-token entropy than in open-ended language generation. The paper adopts DFlash, a diffusion-based parallel drafter, which generates entire candidate blocks in one forward pass, naturally capturing intra-block correlations. A lightweight two-layer draft model (block size = 8) trained on 60k clips achieves an average accepted length of 5.6 tokens, delivering a 4.7× decoding speedup over the unoptimized baseline (2.9× over the system-optimized baseline).

Adaptive-Step Flow Matching (Action): The flow-matching action head converts the VLM's hidden representations into trajectory waypoints through iterative denoising (8 steps in the evaluated configuration). Profiling the velocity field reveals a striking structure: "the normalized relative differences between consecutive velocities trace a U-shape... The velocity changes sharply at the first and last steps, where the trajectory departs from the noise prior and converges onto the data manifold, but is nearly constant through the middle." FlashDrive exploits this by caching the velocity at the middle steps and reusing it, replacing four intermediate evaluations with cached velocities. This cuts action latency from 113.9 ms to 47.6 ms while being near-lossless: minADE6 increases by only 0.04 m, while minADE1 improves by 0.14 m.

Quantization: The paper applies ParoQuant to quantize the VLM backbone's weights to 4-bit and executes inference with 8-bit activations using the W4A8 variant of Marlin kernels. The action expert remains in BF16. This cuts the memory footprint to 18.3 GB and reduces end-to-end latency by a further 14% (176.0 ms → 151.4 ms).

System Optimizations: Each pipeline stage is compiled into a CUDA Graph, and attention and MLP kernels are fused (e.g., Q, K, V projections and gate/up projections in the MLP are fused from six launches to two). Together, CUDA Graphs and kernel fusion provide a 1.40× speedup without changing the model's computation.

The techniques compound: applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717 ms to 151 ms (4.7×), raising the control frequency from 1.4 Hz to 6.6 Hz on a single GPU. The acceleration is nearly lossless: minADE6 @6.4s degrades by only 0.08 m (from 0.767 m to 0.844 m), while minADE1 @6.4s improves from 1.705 m to 1.573 m. The paper suggests streaming fine-tuning acts as a regularizer that reduces prediction variance.

Cross-device deployment results show consistent speedups ranging from 4.0× to 6.0× across the Jetson Thor, RTX 3090, RTX 4090, RTX 5090, and RTX PRO 6000 with a single trajectory sample. With six trajectory samples, FlashDrive reaches 9.6× on Jetson Thor, 10.0× on RTX 5090, and 10.6× on RTX PRO 6000. Notably, on RTX 3090 and RTX 4090 (24 GB VRAM), the unoptimized model fails due to memory limits while FlashDrive still runs.

Closed-loop evaluation in AlpaSim shows FlashDrive preserves driving quality: collision rate drops from 0.19 to 0.15, off-road rate from 0.41 to 0.32, and plan-deviation score falls from 0.24 to 0.16. The one metric that regresses is Wrong Lane (0.45 to 0.51), which is particularly sensitive near intersections. FlashDrive also achieves a 2.5× per-step rollout speedup in AlpaSim (from 1150 ms to 463 ms), improving closed-loop training and evaluation efficiency.

The paper concludes that the path to efficiency is not one universal technique applied everywhere, but the right lightweight shortcut matched to each stage, and that this profile-then-exploit methodology generalizes broadly to any inference pipeline with structurally heterogeneous bottlenecks.

Improvements for AI systems

Improvements to AI Systems:

  1. Stage-Aware Redundancy Exploitation: Implement a profile-then-exploit pipeline that identifies and eliminates redundancy at each inference stage independently—streaming visual encoding with KV-cache reuse across temporal frames, diffusion-based speculative decoding for structured reasoning, and adaptive-step denoising with cached intermediate velocities—rather than applying a single global optimization.

  2. Temporal KV-Cache Persistence with Pre-RoPE Caching: For any multi-frame or multi-view vision-language model, cache keys before rotary position embedding and apply RoPE on-the-fly, enabling seamless insertion of new frames without full re-encoding, reducing visual compute by 75% in continuous operation.

  3. Diffusion-Based Parallel Drafting for Structured Generation: Replace autoregressive token-by-token drafting with a lightweight diffusion model that generates entire candidate blocks in one forward pass, specifically tuned for low-entropy, template-driven reasoning chains (e.g., driving decisions, code generation with fixed structure, form filling), achieving 4.7× decoding speedup.

  4. Velocity-Field Profiling for Adaptive Denoising Steps: In any iterative refinement head (flow matching, diffusion, or score-based), profile the per-step change magnitude and cache/reuse intermediate evaluations where the field is near-constant, reducing steps from 8 to 4 with negligible quality loss.

  5. Streaming Fine-Tuning as Regularization: After applying streaming inference, fine-tune the model on the shifted distribution—this not only recovers accuracy but acts as a regularizer, reducing prediction variance and improving metrics like minADE1 beyond the baseline.

  6. Heterogeneous Precision Allocation: Quantize only the large VLM backbone to W4A8 while keeping task-specific heads (e.g., action experts) in BF16, preserving precision where it matters most and cutting memory footprint by 75% (e.g., from 70 GB to 18.3 GB).

  7. CUDA Graph Compilation with Kernel Fusion: Compile each pipeline stage into CUDA Graphs and fuse attention/MLP projections (e.g., QKV and gate/up projections) to reduce kernel launch overhead, yielding 1.40× speedup with zero algorithmic change.

  8. Cross-Device Adaptive Deployment: Design the system to automatically select optimization levels based on available VRAM—enabling inference on memory-constrained devices (e.g., 24 GB GPUs) where the unoptimized model fails, while scaling to higher throughput on larger hardware.

What the Improved AI System Can Do:

  • Real-Time Autonomous Driving: Operate end-to-end vision-language-action models at 6.6 Hz control frequency on a single GPU (up from 1.4 Hz), enabling real-time closed-loop driving with collision rates reduced by 21% and off-road rates by 22%.

  • Memory-Efficient Deployment: Run 10B-parameter VLA models on consumer GPUs with 24 GB VRAM (e.g., RTX 3090/4090) that previously could not load the model, expanding accessibility to edge hardware like Jetson Thor.

  • High-Throughput Structured Reasoning: Generate driving reasoning chains, structured plans, or templated code at 4.7× faster speeds while maintaining output quality, suitable for latency-critical robotics or interactive agents.

  • Near-Lossless Acceleration: Achieve 4.7–10.6× end-to-end speedups across devices with minimal metric degradation (e.g., minADE6 within 0.08 m), and in some cases improve prediction accuracy (minADE1 improves by 0.13 m) due to regularization from streaming fine-tuning.

  • Scalable Simulation and Training: Accelerate closed-loop rollouts by 2.5×, enabling faster reinforcement learning, data collection, and evaluation in simulated environments.

  • Generalizable Efficiency Methodology: Apply the same stage-wise profiling and targeted shortcut approach to other heterogeneous inference pipelines (e.g., multimodal chatbots, video understanding, robotic control) to achieve similar compound speedups without sacrificing performance.

Abstract

Vision-Language-Action (VLA) models promise to bring end-to-end reasoning to autonomous driving, but their computational cost remains far too high for real-time control. The core challenge is structural: VLA inference is not a single bottleneck but a cascade of four. Visual encoding wastes compute on overlapping video frames; language-model prefill recomputes context that could be carried over from the previous timestep; reasoning tokens are generated serially despite low entropy; and flow-matching denoising applies uniform compute to a non-uniform velocity field. Addressing any one stage in isolation leaves the others untouched. We propose FlashDrive, an algorithm-system co-design framework that targets all four stages simultaneously. Our key insight is that each bottleneck admits a distinct, lightweight algorithmic shortcut: temporal overlap enables streaming KV-cache reuse across frames; the low per-token entropy and strong intra-block correlations of driving-domain reasoning make a non-autoregressive diffusion drafter highly effective for speculative decoding; and the velocity field's structure---sharp at the endpoints, flat in the middle---permits adaptive step caching that concentrates compute where it matters. Layered on system-level CUDA Graph compilation and kernel fusion, these techniques compound. Applied to Alpamayo 1.5-10B with W4A8 quantization, FlashDrive reduces end-to-end latency from 717ms to 151ms (4.7x) while leaving accuracy essentially unchanged: minADE6@6.4s shifts by only 0.08m, minADE1 improves, and closed-loop collision and off-road rates improve in simulation. By raising a 10B-parameter reasoning VLA from 1.4 Hz to 6.6 Hz on a single GPU, FlashDrive moves end-to-end autonomous driving substantially closer to real-time deployment.

Sources

Related papers