SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving".
Dev: The gist: SDPAD, a fully spike-driven end-to-end planning pipeline,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: So we've talked about who wrote this and what the overall goal of "SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving" is. Now let’s look at what they actually propose to solve that problem.
Dev: They introduce three main technical contributions here, so it’s not just one simple idea; it’s a multi-part system designed to handle the whole process of driving decisions from perception through planning.
Taro: I'm paying attention because they focus on how they handle the scene understanding part, specifically mentioning a query-based mechanism for distilling information from dense maps.
Rosa: Right, and one of those parts is this Spike-QFormer, which acts like a unified spiking query transformer to gather context from different sources like the ego vehicle, the surrounding agents, and the map.
Dev: That’s interesting because they say this mechanism fuses that heterogeneous scene information into waypoint predictions using spiking cross-attention, rather than relying on standard floating-point softmax operators.
Taro: The way they are doing that fusion without those floating-point operations sounds like a big step for making these models runnable on neuromorphic hardware.
The paper's summary: Rosa: Building on that, the core idea behind SDPAD is that it replaces traditional continuous methods with spike-driven operations at every stage of the process. It’s about converting everything into integer spike representations in a single feed-forward pass.
Dev: That single pass structure is what really sets it apart from existing SNN planners because they don't have to rely on iterative time-step simulations, which cuts down on latency issues significantly.
Taro: So the system bypasses those slow temporal loops that plague some SNN approaches by doing everything in one go, which makes it much more practical for real-time autonomous driving tasks.
Rosa: It also addresses the 2D to three dee lifting problem with something they call Spike-three dee-Lift, which replaces the floating-point softmax depth distribution with a spike-driven max <ref:2610.11583#pg1>.
Dev: That’s where I see the direct energy saving potential; turning that continuous depth estimation into discrete, integer spike intensities should drastically reduce the computational workload in neuromorphic hardware.
The paper's improvements: Taro: Moving into the specifics of how they make this work better, they propose several key technical innovations, including Spike-three dee-Lift and that query transformer mentioned before <ref:2610.11583#pg1>.
Rosa: They also introduced a specific mechanism called Spike-Driven Depth Lifting or SDM, which quantizes continuous depth estimation into discrete spike states.
Dev: The math behind SDM involves mapping the continuous depth input D to a raw spike activation M using an exponential mapping, and then normalizing it down to a sparse spike output SD using the max function.
Taro: That operation is described as degenerating into a spike-driven linear accumulation in neuromorphic hardware, which sounds like it directly translates to lower FLOPs.
Conclusion: Rosa: So, looking at the whole picture of SDPAD, they are demonstrating that you can get planning accuracy comparable to mainstream ANN planners while using far less energy by leveraging integer spike representations throughout.
Dev: The experimental results on the NAVSIM benchmark show it reaching eighty-six point three PDMS, which is an improvement over previous SNN planners like SAD by about four point three points, and they match those ANN planners at a much lower energy cost for that level of performance.
Taro: I think the ablation studies are telling us that rich interactions between the ego, agent, and map queries are what really drive high-quality planning, because when you combine all three contexts in the Spike-QFormer, you get a way to model those multimodal interactions.
Rosa: And they also showed that increasing their quantization steps from two to eight consistently improves things; for example, the L2 error drops from five point three two m to zero point six four m when using more steps, and the collision rate goes down from six point one two percent to zero point two six percent.
Dev: It’s a very concrete result showing that while they are trading some precision for energy efficiency, the gains in accuracy and safety metrics are substantial enough to matter for deployment.
Taro: So what this means for us is that we're seeing a blueprint where you can do the entire perception-to-planning computation on integer spike representations in one feed-forward pass without needing those iterative simulation loops.
Rosa: That’s the big picture here—a system that promises to be much more efficient for edge AI vehicles by leveraging this whole spike pipeline concept in "SDPAD: A Fully Spike-Driven Pipeline for End-to-End Autonomous Driving."
Chengjun Zhang, Yuhao Zhang, Jie Yang, Mohamad Sawan
Zhejiang Key Laboratory of 3D Micro/Nano Fabrication and Characterization · Integrated-On-Chips Brain-Computer Interfaces Zhejiang Engineering Research Center · CenBRAIN Neurotech
cs.RO, cs.AI
Submitted: 2026-10-08
Updated: 2026-10-08
The gist: The gist: SDPAD, a fully spike-driven end-to-end planning pipeline, achieves planning accuracy comparable to mainstream ANN planners while consuming less than 2% of their energy by converting
Key concepts
- SDPAD Architecture Overview
- This is the end-to-end pipeline that takes pre-trained ANN perception, converts it to integer spikes, lifts 2D images into bird's-eyeview (BEV) using spike-driven max depth distribution, and plans using a SpikeQFormer. It's designed for energy efficiency by operating entirely on spike representations.
- Spike-Driven Max (SDM)
- SDM is a technique that converts continuous depth estimation into discrete, integer-valued spike intensities. It uses an exponential mapping to capture deeper ranges, resulting in a sparse spike output that mimics linear accumulation in neuromorphic hardware, significantly reducing computational load.
- Spike-QFormer
- This is a unified spiking query transformer that takes context from ego (vehicle), agent (other vehicles), and map data. It fuses these contexts into waypoint predictions using spiking cross-attention, allowing the system to generate trajectories based on integrated scene understanding.
Terminology
Summary
The gist: SDPAD, a fully spike-driven end-to-end planning pipeline, achieves planning accuracy comparable to mainstream ANN planners while consuming less than 2% of their energy by converting perception and planning into integer spike representations in a single feed-forward pass.
SDPAD Architecture Overview
SDPAD is introduced as a fully spike-driven pipeline that closes the gap between spiking neural networks (SNNs) and dense artificial neural networks (ANNs) in end-to-end autonomous driving by converting a pre-trained ANN perception stack into integer-spike form via quantized ANN2SNN conversion, lifting multi-view images into the bird’s-eyeview (BEV) space with a spike-driven max (SDM) depth distribution, and planning through the SpikeQFormer which fuses ego, agent, and map queries via learnable waypoint queries through cross-attention followed by deformable spike-cross-attention refinement
Key Technical Innovations
The paper proposes three core technical contributions to SDPAD
-
We propose SDPAD, to our knowledge the first fully spike-driven end-to end planner evaluated in closed loop on NAVSIM
-
We propose Spike-3D-Lift, which replaces the floatingpoint softmax depth distribution with a spike-drivenmax, turning 2D to 3D BEV lifting into hardwarefriendly spike gated accumulation while preserving geometric fidelity
-
We propose the Spike-QFormer, a unified spiking query transformer that distills ego, agent, and map context into dedicated query groups and fuses them into waypoint predictions via spiking cross-attention, together with a deformable spike-cross-attention module for trajectory refinement
Spike Dynamics and Conversion
The system employs the iterative Leaky Integrateand-Fire (I-LIF) model neuron with a soft reset mechanism to define spiking neuron dynamics The integer ANN2SNN conversion is achieved by mapping the layer-l SNN firing rate to the ANN activation via a formula involving a quantized clip activation, which minimizes the conversion gap This integer-driven framework effectively minimizes the mapping error and mitigates residual potential variance
Spike-Driven Depth Lifting (SDM)
To achieve energy efficient 2D to 3D feature lifting, the SpikeDriven-Max (SDM) mechanism quantizes continuous depth estimation into discrete, integer-valued spike intensities Specifically, given the continuous depth input D, it first restricts its range and applies a rounding operation to generate discrete spike states Dˆ To capture the non linear significance of deeper ranges, an exponential mapping is employed to yield the raw spike activation M=Dˆ · 2(Dˆ−Dmax) The final normalized sparse spike output SD is formulated as SD = max(M, ϵ) / N This operation degenerates into a spike-driven linear accumulation in neuromorphic hardware, drastically reducing computational FLOPs while preserving representation capacity
Query-Based Planning Decoder
The Spike-QFormer is a fully spike-driven query transformer that aggregates heterogeneous driving context into waypoint predictions It builds upon a single shared primitive—spiking attention—which is instantiated into two query decoders that progressively distill the BEV scene representation into three complementary query groups (ego, agent, and map), followed by a waypoint query decoder that fuses them for trajectory generation The attention output is then computed without softmax normalization because Q˜, K˜ and V˜ are all integer-valued spike tensors
Deformable Refinement
The waypoint tokens are refined with a Deformable Spike-CrossAttention (Deformable SCA) module which couples the adaptive spatial sampling of deformable attention with an all-spike computation paradigm This module replaces every floatingpoint attention computation with spikedriven operations, where the sampling locations are obtained by combining the reference points with the predicted offsets, normalized by the spatial shape of each feature level
Experimental Validation
Extensive experiments on nuScenes and NAVSIM show that SDPAD substantially narrows the gap between SNN and ANN planners, achieving planning accuracy comparable to mainstream ANN baselines at a small fraction of their energy consumption On the NAVSIM benchmark, SDPAD reaches 86.3 PDMS, surpassing the previous SNN planner SAD by 4.3 points and matching mainstream ANN planners at a fraction of their energy The complete SDPAD leverages the integer spike representation of spiking neurons to represent the dynamic information of the driving scene, demonstrating that a fully spike-driven pipeline can maintain exceptional safety and planning precision in realistic navigation tasks
Ablation Study Insights
The ablation studies confirm that rich query interactions and attention mechanisms are crucial for high-quality planning, as the full combination of Ego, Agent and Map Query yields an L2 of 0.66 m and a collision rate of 0.30% Furthermore, increasing the quantization steps from 2 to 8 consistently improves the planning quality: L2 decreases from 5.32 m to 0.64 m, and the collision rate drops from 6.12% to 0.26% The SDM module reduces computational overhead by 23.6% while maintaining competitive perception performance, achieving an L2 error of 0.64m and a low collision rate of 0.26% at 3s
Conclusion
SDPAD provides a blueprint for future ultra-efficient edge AI vehicles by performing the entire perception-toplanning computation on integer spike representations in a single feed-forward pass, without any temporal simulation loops The paper concludes that SDPAD performs the entire perception-toplanning computation on integer spike representations in a single feed-forward pass, without any temporal simulation loops
References
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh Vora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11621–11631, 2020
Li Chen, Penghao Wu, Kashyap Chitta, Bernhard Jaeger, Andreas Geiger, and Hongyang Li. End-to-end autonomous driving: Challenges and frontiers. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12):10164–10183, 2024a
Shaoyu Chen, Bo Jiang, Hao Gao, Bencheng Liao, Qing Xu, Qian Zhang, Chang Huang, Wenyu Liu, and Xinggang Wang. Vadv2: End-to-end vectorized autonomous driving via probabilistic planning. arXiv preprint arXiv:2402.13243, 2024b
Kashyap Chitta, Aditya Prakash, Bernhard Jaeger, Zehao Yu, Katrin Renz, and Andreas Geiger. Transfuser: Imitation with transformer-based sensor fusion for autonomous driving. IEEE transactions on pattern analysis and machine intelligence, 45(11):12878–12895, 2022
OpenScene Contributors. Openscene: The largest upto-date 3d occupancy prediction benchmark in autonomous driving. In Proceedings of the Conference on Computer Vision and Pattern Recognition, Vancouver, Canada, pages 18–22, 2023
Daniel Dauner, Marcel Hallgarten, Tianyu Li, Xinshuo Weng, Zhiyu Huang, Zetong Yang, Hongyang Li, Igor Gilitschenski, Boris Ivanovic, Marco Pavone, et al. Navsim: Data-driven non-reactive autonomous vehicle simulation and benchmarking. Advances in Neural Information Processing Systems, 37:28706–28719, 2024
Jianhao Ding, Zhaofei Yu, Yonghong Tian, and Tiejun Huang. Optimal ann-snn conversion for fast and accurate inference in deep spiking neural networks.
Improvements for AI systems
-
No temporal simulation loops in inference will be eliminated, as SDPAD is described as performing
inference is a single feed-forward pass without temporal simulation loops,
which directly addresses the latency overhead of existing SNN planners thatrely on iterative time-steps simulations.
-
The system will achieve high energy efficiency by converting the perception stack into integer spike form via
quantized ANN2SNN conversion
and usingspike-driven-max (SDM) depth distribution,
which replaces floating-point softmax distributions, leading to a consumption of69.9 mJ—less than 2% of recent ANN baselines.
-
The planning decoder will utilize the Spike-QFormer, which distills context from ego, agent, and map queries via
spiking decoder layers
and fuses them using asingle-layer spiking cross-attention,
enabling it tomodel the multimodal interactions among the ego vehicle, surrounding agents, and map context.
-
Trajectory refinement will be enhanced by incorporating a
deformable spike-cross-attention module anchored to the predicted trajectory itself,
which leverages adaptive spatial sampling via Deformable SCA to ground waypoint tokens in fine-grained spatial evidence.
Sources
- VADv2: End-to-End Vectorized Autonomous Driving via Probabilistic Planning
- Optimal ANN-SNN Conversion for Fast and Accurate Inference in Deep Spiking Neural Networks
- SAH-Drive: A Scenario-Aware Hybrid Planner for Closed-Loop Vehicle Trajectory Generation
- Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
- Recent Advances and New Frontiers in Spiking Neural Networks
- ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving
- Spikformer: When Spiking Neural Network Meets Transformer
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving