WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA".
Jane: The paper was written by Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He et al. from Fudan University and Yinwang Intelligent Technology Co.,Ltd.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Alright, welcome back to the show, everybody. I'm Tom, and as always, I'm here with my co-host, Jane. Today we are digging into a paper that has a bit of a mouthful for a title: "WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA."
Jane: And Tom, I have to say, that title is dense, but what it represents is actually a pretty wild idea. It's from a team at Fudan University and Yinwang Intelligent Technology, and they're basically asking, "What if we could take a really smart, but slow, AI driver and turn it into a really fast one without losing its brains?"
Tom: Exactly. So, the "AR" in that title stands for autoregressive. That's the tech behind models like ChatGPT, where the AI generates text, or in this case, driving commands, one word at a time, in a sequence. It's smart, but it's slow because it has to wait for each word before it can start the next one.
Jane: And that's a huge problem for a self-driving car. You can't have a car that takes a quarter of a second to decide what to do next when it's barreling down the highway. The paper says the baseline model takes twenty-two point seven milliseconds per token. That might sound fast, but for a car, that's an eternity.
Tom: Right, so they came up with this idea to convert that slow, sequential "thinker" into a "diffusion" model. And diffusion models are the things that generate images. They start with a bunch of noise and slowly refine it into a picture. The team applied that same concept to driving commands.
Jane: So instead of reading a sentence word-by-word, the car can now look at a whole block of "noisy" driving commands and refine them all at once, in parallel. That's the "Diff" part of the title. And this simple switch, just going from sequential to parallel, gives them a two point eight times speedup.
Tom: But here's the catch, and this is what I love about this paper. You can't just flip a switch. The "attention" inside these models works completely differently. An autoregressive model only looks at the past, but a diffusion model needs to look at everything at once. So they had to build this bridge, this "Hierarchical Distillation" process, to transfer the knowledge without breaking the model.
Jane: And that's the real magic here. They're not training a new model from scratch. They're taking a pre-trained, intelligent model and carefully reshaping its brain to think in parallel. It's like taking a brilliant professor who reads books page-by-page and teaching them to speed-read an entire chapter at once, without losing comprehension.
Tom: So we've got the "what" and the "why". The "how" is this three-stage process, and I think that's where we should go next. Because it's not just one trick; it's a whole curriculum.
Jane: Oh, definitely. Let's talk about that three-stage process. It's the core of the paper, and it's really clever engineering.
Summary: Jane: So, Tom, we've established that "WAM-Diff2" is all about taking a smart but slow AI driver and making it fast. The secret sauce is this three-stage training process they call "Hierarchical Distillation." Let's break it down.
Tom: Please, because it sounds like they're just shrinking a model, but it's way more nuanced than that. The first stage is all about baby steps. They call it "Progressive Block-Wise Adaptation." Instead of forcing the model to think in parallel all at once, they start by letting it look at just a few tokens at a time.
Jane: Right. They start with a block size of one, which is the old, slow way. Then they bump it up to four, then eight, then sixteen, and finally thirty-two. Each time, they're giving the model a little more freedom to look around and think in parallel. It's a curriculum for the AI.
Tom: And this is where it gets interesting. The second stage is where they actually teach it to be a good diffusion model. They call it "Block-Wise Distillation." They take the stable, small-block model from stage one and use it as a teacher for the larger-block model. The student model is trying to predict the same answers as the teacher, but on noisy, corrupted inputs.
Jane: And that's the key to fixing "exposure bias." In the old autoregressive world, the model only ever saw correct, ground-truth data during training. But during real driving, it makes mistakes, and those mistakes compound. It's like learning to ride a bike with training wheels and then being thrown onto a rocky trail. The diffusion model, on the other hand, is trained on messy, partially-correct data, so it learns to recover from its own errors.
Tom: So it's not just about speed; it's about robustness. And the third stage, "Model-Wise Cross-Scale Distillation," is about bringing in the big guns. They take an eight-billion parameter diffusion model and use it to teach a much smaller, two-billion parameter student.
Jane: And that's the part that blew my mind. The paper has this metric called "top-K overlap." It measures how often the teacher and student agree on the most likely next token. And they found that two diffusion models agree with each other much more than an autoregressive model and a diffusion model do. They share the same "thinking pattern."
Tom: So the smaller model can learn from the bigger one much more effectively because they're speaking the same language. It's like a chess grandmaster teaching a student, but they both play the same opening. The lessons just make more sense.
Jane: And the results are pretty stunning. They took this two-billion parameter model and got it to match the performance of the eight-billion parameter autoregressive baseline on driving tests, but at a fraction of the latency. On the NAVSIM benchmark, they went from eighty-eight point one four PDMS to eighty-seven point four four, which is basically a tie, but they're doing it two point eight times faster.
Tom: And that's just the algorithmic speedup. They also threw in some serious systems engineering with FlashInfer and CUDA Graphs, which we should talk about, because that's where the fifteen point one times total speedup comes from.
Jane: Oh, absolutely. We can't just talk about the AI magic; we have to talk about the engineering that makes it actually run on a car.
Improvements: Tom: So, Jane, we've covered the "what" and the "how," but now I want to get into the nitty-gritty of the "how fast." The paper claims a cumulative fifteen point one times speedup, and that's not just from the diffusion trick. That's from some serious systems-level optimization.
Jane: Right, and this is where I think a lot of people get lost. The two point eight times speedup is from the algorithm itself. But to get to fifteen point one times, they had to make the hardware work harder for them. They used something called FlashInfer, which is a super-efficient way to run the attention mechanism on a GPU.
Tom: And then they used CUDA Graphs. That's a way to pre-record a sequence of GPU operations so the computer doesn't have to re-launch them every single time. It's like having a pre-planned route for a delivery driver instead of making them check a map at every intersection. That alone gave them another three point one times boost.
Jane: And the combined effect is that they went from twenty-two point seven milliseconds per token down to just one point five milliseconds. That's a fifteen-fold reduction in latency. For a self-driving car, that's the difference between reacting to a pedestrian and not reacting in time.
Tom: But here's the thing I find most impressive. They didn't just make it faster; they made it more reliable over time. The paper has this great figure showing the "per-waypoint L2 error." That's basically how far off the car's predicted path is from the actual path. And in the autoregressive model, that error grows and grows as the prediction gets longer.
Jane: It's like a game of telephone. The first word is right, but by the eighth word, the message is garbled. The diffusion model, because it can look at the whole path at once, doesn't have that problem. The error stays low and flat. It's not just faster; it's fundamentally more stable.
Tom: And that's a huge deal for real-world driving. You need a system that can predict a path four seconds out and have it still be relevant. The paper shows that the diffusion model's error at the final waypoint is significantly lower than the autoregressive model's.
Jane: And they tested this on a bunch of benchmarks, not just one. They looked at open-loop planning with NAVSIM, closed-loop driving in a simulator called Bench2Drive, and even general visual understanding and object detection. The model holds its own across the board.
Tom: So we have a model that's faster, more robust, and can still do all the other stuff a driving AI needs to do, like understand the scene and detect objects. It's a pretty compelling package.
Jane: It really is. And it makes you wonder, if this approach works for driving, where else could it be applied? That's the big-picture question we should tackle as we wrap up.
Conclusion: Tom: Alright, Jane, let's bring it home. We've spent this whole episode on "WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA," and I think we've only scratched the surface of its potential.
Jane: I agree. The core idea is so elegant. They took the intelligence of a large, slow model and, through this three-stage distillation process, transferred it into a small, fast, and robust diffusion model. It's a blueprint for making any AI system deployable in the real world.
Tom: And the numbers speak for themselves. A fifteen point one times speedup is not an incremental improvement; it's a paradigm shift. We're talking about taking a system that was too slow for real-time driving and making it fast enough to not just be a passenger, but to be the driver.
Jane: And the fact that it also fixes exposure bias is the cherry on top. It's not just about speed; it's about trust. You need an AI driver that doesn't make compounding mistakes, and this paper shows a clear path toward that.
Tom: So, as we say goodbye to this paper, I think the biggest takeaway is that the future of AI isn't just about bigger models. It's about smarter ways to use the models we already have. This hierarchical distillation approach is a powerful tool for that.
Jane: It really makes you wonder where this could go next. If you can translate a language model into a diffusion model, can you do the same for other tasks? Robotics, video generation, even drug discovery? The possibilities are huge.
Tom: Well, that's all the time we have for today. Thanks for joining us on this deep dive into "WAM-Diff2." It's been a fascinating look at the future of autonomous driving and AI efficiency.
Jane: And we're already looking forward to the next paper. There's always something new and exciting on the arXiv. Until next time, keep your eyes on the road and your mind on the future.
Tom: See you all next time.
Zhihao Zhu, Hanlin Shang, Mingwang Xu, Feipeng Cai, Zhuolin He, Yaoyi Li, Jianhua Han, Hang Xu, Siyu Zhu
Fudan University · Yinwang Intelligent Technology Co.,Ltd
cs.RO, cs.AI, cs.CV
Submitted: 2026-08-18
Updated: 2026-08-19
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 69/100
The gist: "We introduce WAM-Diff2, a multi-task discrete diffusion VLA framework powered by a novel three-stage hierarchical distillation strategy.
Key concepts
- AR (Autoregressive)
- This is the technology behind models like ChatGPT where the AI generates text sequentially, one word after another. It is smart but slow because it must wait for each previous word before generating the next one.
- Diffusion Model
- These are models that generate output by starting with noise and slowly refining it into a final result. The paper applies this concept to driving commands, allowing the car to refine all commands in parallel instead of sequentially.
- Hierarchical Distillation
- This is a three-stage training process used to transfer knowledge from a slow, pre-trained model to a smaller, faster diffusion model. It involves progressive block adaptation and cross-scale distillation to reshape the model's thinking pattern.
- Exposure Bias
- This is an issue where models trained only on correct data make mistakes when faced with real-world, noisy inputs. Diffusion models are trained on messy data, which helps them learn to recover from these errors during actual driving.
Terminology
Summary
Summary
The paper introduces WAM-Diff2, a multi-task discrete diffusion Vision-Language-Action (VLA) framework for autonomous driving, designed to transform pre-trained autoregressive generalists into highly efficient parallel diffusion models. The authors state: "We introduce WAM-Diff2, a multi-task discrete diffusion VLA framework powered by a novel three-stage hierarchical distillation strategy. Instead of initiating training from scratch, our framework seamlessly transforms any advanced, pre-trained autoregressive driving generalist into an ultra-efficient parallel refinement architecture."
The core problem addressed is the dichotomy between autoregressive generalists, which offer broad multi-task semantic reasoning but suffer from high computational latency and exposure bias due to sequential decoding, and specialized diffusion policies, which enable low-latency parallel execution but are typically narrow, single-task systems lacking holistic visual-linguistic reasoning. The authors argue: existing paradigms present a rigid compromise between multi-task cognitive intelligence and parallelized execution efficiency.
Their solution is a universal architecture translation paradigm
that bridges the gap between causal autoregressive attention and bidirectional parallel refinement.
The methodology is built upon a three-stage hierarchical distillation strategy. Stage I (Progressive Block-Wise Adaptation) introduces a block-causal attention mechanism that incrementally relaxes causal constraints through a curriculum-based expansion of decoding blocks (B = 1 → 32), ensuring mathematical stability during the attention shift. Stage II (Block-Wise Distillation) uses a stable, small-block diffusion teacher to guide larger-block student variants, employing a symmetric Jensen-Shannon Divergence (JSD) loss over intermediate noisy states to enforce paradigm consistency and eliminate exposure bias. Stage III (Model-Wise Cross-Scale Distillation) transfers holistic semantic capabilities from an advanced 8B diffusion teacher to a highly efficient 2B student, leveraging the discovery that models sharing the same diffusion paradigm exhibit highly aligned token prediction patterns.
The architecture builds upon the Qwen3-VL framework, with two variants: an 8B variant integrating a SigLIP2-SO-400M visual encoder with 36 Transformer blocks, and a 2B variant using a SigLIP2-Large encoder with 28 Transformer blocks. All modalities—linguistic tokens, 2D bounding boxes, and future waypoints—are processed through a unified text tokenizer, removing task-specific projection heads. The standard causal attention mask is modified into a block-causal pattern where tokens within the same decoding block employ bidirectional parallel attention while inter-block dependencies retain causal constraints.
The discrete diffusion formulation uses an absorbing-state (masked) approach. The forward process progressively corrupts clean token sequences into noisy states over T discrete timesteps, with tokens either retaining identity or transitioning to a special [MASK] token. During inference, the model predicts all token distributions simultaneously, fixing a subset of tokens with highest prediction confidence at each denoising step and re-masking the remaining tokens, compressing full-sequence decoding complexity to O(T) steps where T ≪ L.
System-level optimizations include FlashInfer, which provides customized attention kernels optimized for the block-causal attention pattern, and CUDA Graphs, which encapsulate the entire execution graph to eliminate CPU launch overheads. These optimizations yield a cumulative 15.1× reduction in decoding latency, from 22.7 ms/token to 1.5 ms/token.
The training pipeline spans three phases. Phase 1 involves multi-task autoregressive pretraining for 5 epochs using standard causal next-token prediction. Phase 2 involves progressive adaptation and block-wise distillation, gradually relaxing causal constraints by expanding block size B ∈ 4, 8, 16, 32, with each variant initialized from its predecessor's weights. Phase 3 involves model-wise cross-scale distillation, where the 2B student is distilled for 5 epochs against the advanced 8B diffusion teacher using JSD loss.
Evaluation covers five benchmarks: NAVSIM (v1/v2) for open-loop trajectory planning, Bench2Drive for closed-loop planning, DriveBench and LingoQA for driving-oriented visual question answering, and COCO for visual perception. The authors report that WAM-Diff2 achieves strict performance parity or superiority relative to its autoregressive foundation while delivering an obvious inference speedup.
Key quantitative results include: the 2B diffusion model with B = 32 achieves a DriveBench score of 48.80 and a LingoQA metric of 65.80, matching or exceeding mature autoregressive generalists. For spatial perception, it retains 36.3 mAP on COCO. On NAVSIM v1, it delivers 87.44 PDMS without auxiliary reinforcement fine-tuning or score-based candidate selection. With score-based candidate selection, it reaches 91.1 PDMS on NAVSIM v1 and 90.7 EPDMS on NAVSIM v2, outperforming specialized agents like ReCogDrive and DriveVLA-W0. On Bench2Drive, it attains a Success Rate of 49.55% and a Driving Score of 78.93, with pronounced improvements in safety-critical scenarios like Emergency Braking (61.67%) and Traffic Sign compliance (85.26%).
The ablation studies demonstrate the effectiveness of each stage. Direct AR-to-diffusion adaptation results in a significant performance drop to 84.1 PDMS, while block-wise distillation recovers performance to 87.7 PDMS, and model-wise distillation further closes the gap to 88.3 PDMS. The authors also show that the 8B diffusion teacher exhibits significantly higher top-K token overlap (ρ5 = 58.2%) with the 2B diffusion student compared to the 8B autoregressive teacher (ρ5 = 51.2%), and that distilling from the autoregressive teacher causes a severe performance drop to 79.9 PDMS. The symmetric JSD loss achieves peak performance (88.3 PDMS) compared to forward KL (88.2) and reverse KL (88.0).
The paper also demonstrates mitigation of exposure bias through bidirectional attention and iterative refinement. On 12,146 paired NAVSIM samples, the block-wise diffusion framework reduces average per-waypoint L2 error by 5.8% (from 0.5935 to 0.5589), with the absolute error reduction widening monotonically over time—from 0.002 at the initial waypoint to 0.082 at the final horizon (waypoint 8)—demonstrating that bidirectional parallel refinement effectively mitigates long-horizon error accumulation.
The authors acknowledge two primary limitations: the structural reliance on discrete tokenization can introduce spatial quantization artifacts affecting trajectory smoothness, and the downstream capabilities of the parallel diffusion student remain fundamentally upper-bounded by the baseline semantic reasoning proficiency of the initial autoregressive teacher. The paper concludes that WAM-Diff2 provides a scalable, low-cost pipeline to deploy advanced multimodal driving intelligence under the rigid throughput constraints of highly efficient autonomous systems.
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can implement in AI systems, followed by the enhanced capabilities of the improved system.
-
Stage I – Progressive Block-Wise Adaptation: Replace the standard causal attention mask with a block-causal mask. Start with block size B=1 (autoregressive) and incrementally expand to B=4, 8, 16, 32. Initialize each larger-block model from the weights of its smaller-block predecessor to ensure stable attention transition.
-
Stage II – Block-Wise Distillation: Use a stable small-block diffusion teacher (e.g., B=4) to distill larger-block students (B=8, 16, 32). Apply a symmetric Jensen-Shannon Divergence (JSD) loss on intermediate noisy states, rather than forward or reverse KL, to preserve multi-modal trajectory distributions.
-
Stage III – Model-Wise Cross-Scale Distillation: Transfer knowledge from an 8B diffusion teacher to a 2B diffusion student. This is effective because diffusion models share higher top-K token overlap (ρ5=58.2%) compared to autoregressive teachers (ρ5=51.2%), enabling efficient capacity compression without losing reasoning ability.
-
Replace causal attention with a block-causal pattern: tokens within the same decoding block use bidirectional parallel attention, while cross-block dependencies remain causal. This allows simultaneous token refinement within each block, reducing decoding latency from O(L) to O(T) where T ≪ L.
-
Integrate FlashInfer kernels optimized for block-causal attention to maximize shared memory throughput during intra-block refinement.
-
Use CUDA Graphs to encapsulate the entire denoising execution graph, eliminating CPU launch overhead for static tensor shapes.
-
Use a masked diffusion process where tokens transition to a special [MASK] token with probability βt. Train the network to reconstruct clean sequences from corrupted observations via random state sampling.
-
During inference, use parallel remasking: predict all token distributions simultaneously, fix the highest-confidence tokens, and re-mask the rest across T denoising steps (typically 8–16 steps for saturation).
-
Instead of training only on ground-truth prefixes (teacher forcing), expose the student to intermediate noisy states generated during parallel decoding. This aligns the training distribution with inference-time conditions, preventing compounding drift errors in long-horizon trajectory prediction.
-
Represent all outputs—natural language, bounding boxes, and waypoints—as text tokens using a shared tokenizer. Remove task-specific projection heads. This enables a single model checkpoint to handle driving understanding, perception, and planning without architectural changes.
-
Decoding latency drops from 22.7 ms/token to 1.5 ms/token (cumulative speedup from AR baseline). Throughput reaches 673.4 tokens/second with hardware optimizations, enabling real-time decision-making for safety-critical driving scenarios.
-
On NAVSIM v1, achieves 88.3 PDMS (vs. 88.1 for AR baseline); on NAVSIM v2, achieves 88.6 EPDMS. On DriveBench, scores 48.80 (vs. 51.23 for AR baseline, a minor 4.7% drop); on COCO mAP, scores 36.3 (vs. 39.2). This demonstrates that the diffusion paradigm preserves holistic cognitive reasoning while enabling parallel execution.
-
Reduces average per-waypoint L2 error by 5.8% (from 0.5935 to 0.5589) on NAVSIM. Error reduction widens monotonically over time—from 0.002 at waypoint 1 to 0.082 at waypoint 8—proving that bidirectional refinement prevents compounding drift.
-
By adjusting block size B from 4 to 32, the system can trade between 68.3 and 124.8 TPS (vanilla) or 401.4 to 673.4 TPS (optimized) with negligible planning degradation (∆ ≤ 0.5 PDMS). This allows deployment across different hardware constraints without retraining.
-
On Bench2Drive, achieves 49.55% Success Rate and 78.93 Driving Score. Notably, excels in safety-critical conditions: 61.67% success in Emergency Braking and 85.26% in Traffic Sign compliance, outperforming many specialized planners.
-
The 2B student inherits complex reasoning from an 8B teacher (via Stage III distillation) without sacrificing parallel decoding speed. This enables deployment of advanced driving intelligence on resource-constrained edge devices (e.g., in-vehicle NPUs).
-
The framework can transform any pre-trained autoregressive VLA (e.g., Qwen3-VL, InternVL3) into a diffusion agent, bypassing the need to train specialized diffusion policies from scratch. This reduces training cost and leverages existing large-scale semantic knowledge.
Sources
- On-Policy Distillation of Language Models for Autonomous Vehicle Motion Planning
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
- The pitfalls of next-token prediction
- Qwen3-VL Technical Report
- Qwen2.5-VL Technical Report
- Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens
- DFM-VLA: Iterative Action Refinement for Robot Manipulation via Discrete Flow Matching
- TransFuser: Imitation with Transformer-Based Sensor Fusion for Autonomous Driving
- ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
- DriveFine: Refining-Augmented Masked Diffusion VLA for Precise and Robust Driving
- ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation
- PyGraph: Robust Compiler Support for CUDA Graphs in PyTorch
- DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
- MiniLLM: On-Policy Distillation of Large Language Models
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- Distilling the Knowledge in a Neural Network
- EMMA: End-to-End Multimodal Model for Autonomous Driving
- Senna: Bridging Large Vision-Language Models and End-to-End Autonomous Driving
- SafeDrive: Fine-Grained Safety Reasoning for End-to-End Driving in a Sparse World
- LLaVA-OneVision: Easy Visual Task Transfer
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving