MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving".
Jane: MomADv2 is a reliable state-space memory framework designed for long-horizon end-to-end autonomous driving, addressing the critical issue where existing temporal memory methods become invalid or misleading during command changes.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Alright, let's talk about who wrote this and what exactly this paper is calling MomADv2. It’s by Song et al., and they have put forward a framework called MomADv2, which is designed to be a more reliable state-space memory system for end-to-end autonomous driving.
Jane: That title really highlights the focus; it’s not just about planning, it's about reliability over long horizons in autonomous driving. It’s aiming to solve that fundamental issue where existing temporal memory methods become invalid or misleading when you make a command change mid-drive.
Lu: The authors are tackling the challenge of selectively leveraging useful historical information while simultaneously suppressing that command-inconsistent memory, which they suggest is a major hurdle for any robust long-horizon system.
Meng: I'm curious about what this means in terms of implementation complexity; building a selective state-space module sounds like it could add significant overhead compared to simpler memory architectures we've used before.
Lalam: From my perspective, the authors are essentially creating a mechanism that learns when historical data is actually useful for the *current* intent, which is a really important concept for any sequential decision-making system.
The paper's summary: Tom: So, looking at the summary of MomADv2: they introduce the Selective State-Space Planning Memory Query Module, or SSM-Q, which filters historical planning queries based on temporal continuity and command consistency to model how planning intentions evolve.
Jane: That sounds like a clever way to manage memory; instead of just storing everything chronologically, they are actively pruning the history that doesn't match what the car is supposed to be doing right now. They store raw queries and baseline commands in a token-keyed memory bank indexed by frame-level identifiers to maintain temporal validity.
Lu: What’s fascinating about their approach is how they retrieve these historical states by first finding command-relevant candidates, and then identifying the "most trajectory-consistent historical candidate under the same command" by minimizing a "temporally shifted trajectory distance."
Meng: That alignment process sounds crucial; if you can compensate for frame-to-frame trajectory offsets, it means you're not just grabbing random old data but finding the most contextually relevant past situation.
Lalam: I think that focus on aligning trajectories ensures that the historical queries being aggregated actually make sense in the current driving context, which is a key improvement over just using raw temporal sequences.
The paper's improvements: Tom: The paper points out two major components they designed: first, this Selective State-Space Planning Memory Query Module (SSM-Q), and second, the Flow-Matching Trajectory Residual Refiner, or FM-Ref.
Jane: The SSM-Q handles the filtering and modeling of intentions, while the FM-Ref steps in later to refine local trajectory deviations by learning a continuous residual correction field from the refined planning output toward the expert trajectory.
Lu: The reliability gate within the SSM-Q is particularly interesting; it computes a reliability indicator based on needing a minimum number of valid states and a distance threshold, ensuring that historical memory is only activated when there's sufficient valid and trajectory-consistent information available.
Meng: That sounds like it prevents the system from relying on too little or too much questionable history, which addresses the practical concern of making decisions with insufficient context.
Lalam: And the FM-Ref’s use of a learnable bounded residual gate to apply corrections judiciously, while preserving the stability of anchor-based planning, shows a sophisticated balance between making improvements and keeping things stable.
Conclusion: Tom: So, wrapping up this discussion on MomADv2: essentially, the framework successfully proposes a reliable state-space memory system that selectively keeps useful history while cutting out inconsistent memory interference through SSM-Q, and it adds FM-Ref to fine-tune local trajectories.
Jane: What this means in practice is that we get consistent improvements in long-horizon planning accuracy and temporal consistency across planning cycles, which leads to better driving safety across various benchmarks like NAVSIM and Bench2Drive.
Lu: The results they show on nuScenes, achieving an average L2 error of one point two one meters and a collision rate of zero point seven six percent over six-second planning compared to MomAD, really validates that selective temporal memory works as intended for safety metrics.
Meng: The validation on closed-loop benchmarks like Bench2Drive reaching a Driving Score of seventy-eight point eight two and the best Comfortness score of thirty-one point two three is what tells us this isn't just theoretical; it shows better real-world interaction capabilities.
Lalam: I feel that this whole approach to reliable state-space memory is going to be important because it teaches sequential decision-making systems how to manage uncertainty in a way that preserves high performance across long sequences.
Tom: Exactly. The final results on NAVSIMv2 navtest showing the lowest Temporal Prediction Consistency errors of one point zero four meters and one point three two meters at the four-second and five-second horizons confirms its superior temporal stability, so this MomADv2 framework is essential for robust end-to-end autonomous driving systems.
Ziying Song, Shengkai Zhang, Lin Liu, Peiliang Wu, Lei Yang*, Dongyang Xu, Bin Sun, Li Wang*, Shaoqing Xu, Caiyan Jia2, Yadan Luo8
School of Artificial Intelligence (School of Software), Yanshan University · Beijing Jiaotong University · Nanyang Technological University · Tsinghua University · China Automotive Technology and Research Center Co., Ltd. · School of Mechanical Engineering, Beijing Institute of Technology · University of Macau · The University of Queensland
cs.CV, cs.RO
Submitted: 2026-08-24
Updated: 2026-09-30
Comments: 16 pages, 6 figures
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 90/100
The gist: MomADv2 is a reliable state-space memory framework designed for long-horizon end-to-end autonomous driving, addressing the critical issue where existing temporal memory methods become invalid or
Key concepts
- Selective State-Space Planning Memory Query Module (SSM-Q)
- This module filters historical planning queries based on temporal continuity and command consistency. It ensures that only relevant past information is retrieved, helping the system model the evolution of planning intentions by suppressing memory that contradicts the current driving command.
- Flow-Matching Trajectory Residual Refiner (FM-Ref)
- This component learns a continuous correction field from the planning output to refine local trajectories against expert paths. It predicts a conditional residual velocity field to correct small deviations, ensuring accurate trajectory execution while maintaining stability through a learnable bounded residual gate.
- Temporal Continuity Token
- This parameter is used within the memory aggregation process to weight historical queries. It helps determine the relevance and reliability of past states by factoring in how continuous or consistent the temporal sequence of planning intentions is, aiding in better filtering.
- Reliability Gate
- This mechanism computes a reliability indicator for historical memory activation. It ensures that past information is only used if a minimum threshold of valid and trajectory-consistent states are available, preventing the system from relying on insufficient or misleading historical data.
Terminology
Summary
MomADv2 is a reliable state-space memory framework designed for long-horizon end-to-end autonomous driving, addressing the critical issue where existing temporal memory methods become invalid or misleading during command changes. This work introduces MomADv2 to selectively leverage useful historical information while suppressing command-inconsistent memory, thereby improving planning continuity and trajectory stability in complex scenarios.
Core Framework and Selective Memory
MomADv2 is built around a selective state-space memory mechanism, specifically the Selective State-Space Planning Memory Query Module (SSM-Q). This module filters historical planning queries based on temporal continuity and command consistency to model the evolution of planning intentions. The framework stores raw planning queries and baseline commands in a token-keyed memory bank, where each entry is indexed by a frame-level identifier that ensures temporal validity. Key phrases include: selectively retains reliable history, rather than fully trusting historical information as MomAD does,
and the mechanism restricts retrieval to candidate subsets associated with the current command.
Reliable Historical Query Retrieval
The SSM-Q employs a multi-step process to retrieve relevant historical states. First, it identifies command-relevant candidates; second, it finds the most trajectory-consistent historical candidate under the same command
by minimizing a temporally shifted trajectory distance.
This alignment compensates for frame-to-frame trajectory offsets, ensuring that retrieved queries are contextually relevant. The retrieval process is further refined by aggregating aligned historical queries using a weighting scheme based on validity and temporal decay, controlled by parameters like temporal continuity token
and decay factors.
Reliability-aware Temporal Modeling
The retrieved historical queries are aggregated into a sequence to model the temporal transition from past to current planning intentions. This sequence is processed by a selective state-space query encoder that models these transitions while preserving candidate-wise planning semantics.
A crucial component is the reliability gate, which computes a reliability indicator based on the minimum number of valid states required and a distance threshold, ensuring that historical memory is activated only when sufficient valid and trajectory-consistent historical information is available.
Flow-Matching Trajectory Residual Refiner
To mitigate local trajectory deviations and error accumulation, MomADv2 incorporates the Flow-Matching Trajectory Residual Refiner (FM-Ref). This module learns a continuous residual correction field from the refined planning output to the expert trajectory. It predicts a conditional residual velocity field
conditioned on both the enhanced planning query and the geometric state of the current trajectory. The refinement process uses this field with an Euler integration scheme, and a learnable bounded residual gate
ensures that corrections are applied judiciously to improve accuracy while preserving the stability of anchor-based planning.
Experimental Validation
Extensive experiments on NAVSIM, Bench2Drive, and nuScenes demonstrate the efficacy of MomADv2. On nuScenes, it achieves an average L2 error of 1.21 m and a collision rate of 0.76% over 6-second planning compared to MomAD. In closed-loop benchmarks like Bench2Drive, MomADv2 reaches a Driving Score (DS) of 78.82 and the best Comfortness score of 31.23, improving upon MomAD's performance by demonstrating more reliable driving via effective temporal memory modeling.
Ablation studies confirm that both SSM-Q and FM-Ref contribute positively to metrics like PDMS, showing that selective temporal memory and flow-matching residual refinement provide complementary benefits for accurate and safe long-horizon planning.
The final results show MomADv2 achieves the lowest Temporal Prediction Consistency (TPC) errors of 1.04 m and 1.32 m at the 4-second and 5-second horizons on NAVSIMv2 navtest, confirming its superior temporal stability.
Conclusion
MomADv2 successfully proposes a framework that selectively preserves useful history while suppressing stale or command-inconsistent memory through SSM-Q, and further refines local trajectories via FM-Ref. This approach yields consistent improvements in long-horizon planning accuracy, temporal consistency across planning cycles, and overall driving safety across various benchmarks. The work confirms that this reliable state-space memory framework
is essential for robust end-to-end autonomous driving systems.
Contributions
The main contributions include:
-
Proposing MomADv2, a reliable state-space memory framework that selectively leverages useful historical information while suppressing invalid memory interference.
-
Designing the Selective State-Space Planning Memory Query Module (SSM-Q) for reliable historical query filtering and intention modeling, and the FlowMatching Trajectory Residual Refiner (FM-Ref) for fine-grained trajectory correction.
-
Demonstrating that MomADv2 effectively improves long-horizon planning consistency and trajectory accuracy on NAVSIM, Bench2Drive, and nuScenes datasets.
-
Showing that the full selective memory strategy achieves the lowest collision rate of 0.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the provided paper, MomADv2: Reliable Temporal Memory for End-to-End Autonomous Driving.
The core innovation lies in replacing indiscriminate historical memory reuse with a selective state-space mechanism (SSM-Q) and refining planning outputs using continuous flow matching (FM-Ref).
Here are specific improvements that can be implemented in AI systems based on the MomADv2 framework, and what these improved systems can achieve:
)
-
A robust, reliable long-horizon planning module capable of maintaining temporal consistency across extended driving horizons without being misled by outdated or command-inconsistent historical data.
-
The ability to dynamically filter historical planning queries based on high-level driving commands (e.g.,
left,
straight
) and temporal continuity tokens, ensuring only relevant and valid past intentions inform current decisions. -
A mechanism to suppress memory interference caused by command shifts or scene changes, leading to significantly reduced trajectory jitter and improved stability in complex urban environments compared to baseline methods like MomAD.
-
Fine-grained local trajectory refinement capabilities that correct small deviations from the expert/anchor trajectory using a learned, query-conditioned residual velocity field (Flow Matching Trajectory Residual Refiner), preserving the stability of anchor-based planning while enhancing smoothness and accuracy.
-
Enhanced safety performance in long-horizon scenarios, demonstrated by a reduction in collision rates (e.g., 15.6% improvement over MomAD under 6-second planning).
-
Superior closed-loop driving performance metrics (Driving Score, Success Rate) on complex interactive benchmarks like Bench2Drive and simulation environments like NAVSIMv2 navhard, indicating better real-world interaction capabilities.
-
Improved trajectory prediction consistency (TPC metric), specifically at longer planning horizons (4s–6s), which translates to more predictable and stable future driving behaviors for the vehicle.
-
Increased performance across a wider range of complex driving abilities, including merging, overtaking, emergency braking, give-way maneuvers, and traffic sign compliance in multi-ability evaluations.
This improved AI system (MomADv2) can perform the following specific tasks:
-
Generate safer and more stable ego trajectories during long-horizon planning (up to 6 seconds), especially when driving commands change abruptly or the environment is dynamic.
-
Maintain consistent driving intentions over extended planning cycles, preventing
command confusion
errors where historical data contradicts the current desired maneuver. -
Achieve high-fidelity trajectory tracking by actively correcting local path deviations using a learned residual field, resulting in smoother and more accurate vehicle motion compared to standard planners that only output initial candidates.
-
Excel in complex, multi-faceted driving tasks like merging into traffic, navigating intersections with dynamic agents, and reacting appropriately to traffic rules (e.g., obeying traffic lights or giving way), as evidenced by its high scores on Bench2Drive metrics.
-
Perform reliably in closed-loop scenarios (simulated real-world interaction) where the vehicle's actions directly influence subsequent sensor observations, leading to better overall driving scores and success rates.
Sources
- Bench2Drive: Towards Multi-Ability Benchmarking of Closed-Loop End-To-End Autonomous Driving
- Hydra-MDP++: Advancing End-to-End Driving via Expert-Guided Hydra-Distillation
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
- Fully Unified Motion Planning for End-to-End Autonomous Driving
- GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- DriveWorld-VLA: Unified Latent-Space World Modeling with Vision-Language-Action for Autonomous Driving
- ME$^3$-BEV: Mamba-Enhanced Deep Reinforcement Learning for End-to-End Autonomous Driving with BEV-Perception
- GraphWorld: Long-Horizon Planning with World Models for End-to-End Autonomous Driving
- SparseDrive: End-to-End Autonomous Driving via Sparse Scene Representation
- GMF-Drive: Gated Mamba Fusion with Spatial-Aware BEV Representation for End-to-End Autonomous Driving
- DRAMA: An Efficient End-to-end Motion Planner for Autonomous Driving with Mamba
- GenAD: Generative End-to-End Autonomous Driving
- DiffusionDriveV2: Reinforcement Learning-Constrained Truncated Diffusion Modeling in End-to-End Autonomous Driving
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models