SimWAM: A Simple World Action Model for End-to-End Autonomous Driving

arXiv:2608.07468 · cs.CV · Submitted 2026-08-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "SimWAM: A Simple World Action Model for End-to-End Autonomous Driving".

Tom: SimWAM presents a novel World-Action Model (WAM) designed to improve end-to-end autonomous driving by leveraging future video prediction as a training signal, thereby avoiding costly test-time future imagination.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, Jane, we’re diving into SimWAM today, a paper that’s been making some interesting moves in autonomous driving research. It tackles a problem where these world models usually get bogged down by having to imagine what happens in the future during the actual driving process.

Jane: That sounds like a really practical issue for real-time systems, Tom. Basically, they're looking at how to make these models better at predicting actions without adding extra computational steps that slow things down significantly when you need immediate decisions.

Lu: What’s fascinating about this paper is how they manage to separate the video prediction part from the action prediction part using a joint flow matching objective. It suggests a way to get those valuable dynamics priors without needing an explicit future-frame generation step, which is where most other methods struggle.

Meng: From my side, I’m curious about the engineering reality of that decoupling. How does this isolation work in practice when you're trying to build something that runs on hardware with strict latency requirements?

Lalam: As an AI, I see this as a cultural shift because it means we can move away from these heavy synthesis steps and focus more on real-time decision making, which fundamentally changes how we trust autonomous systems in everyday life.

Tom: Exactly! The core idea of SimWAM is that instead of asking the planner to generate explicit future visuals, they use the video prediction task during training to teach a lightweight action expert directly about the traffic dynamics.

Jane: That makes sense when you think about it as learning from examples rather than synthesizing new images on demand. They co-train this video expert and action expert using flow matching, which is a clever way to link those two components together during the learning phase.

Lu: The methodology relies on defining a joint training objective, L = L act FM + lambda L vid FM, where both parts of the flow matching equation are instantiated over the action trajectory and future-frame latents, balancing those two signals through that lambda parameter.

Meng: Balancing those two objectives sounds tricky; I wonder how they found that sweet spot for lambda to ensure both experts learn effectively without one overpowering the other during training.

Lalam: It suggests a really elegant way to transfer knowledge—the video expert provides the scene understanding, and the action expert learns how to act based on that understanding, which is a powerful way for an AI system to internalize complex driving behaviors.

Title and authors: Tom: And that leads us into the main advantage they highlight: inference. They use an isolated attention mask so the action prediction stays independent of future frames, meaning no explicit future-frame generation is needed during real-time operation.

Jane: So, if I understand correctly, this means at runtime, the action expert just takes what it sees now and predicts the trajectory based on that representation without having to run a whole visual prediction pipeline first?

Lu: Precisely; they can achieve direct trajectory prediction from the current inputs by conditioning the action expert on z(o t), which is a representation derived from observation at time t, rather than needing explicit future-frame generation during inference.

Meng: That sounds like a massive win for latency, but I have to ask about the isolation itself. What exactly does this isolated attention mask do structurally to keep those two parts truly independent?

Lalam: It really shows how an AI can be modular; you can swap out the video backbone entirely, and the action expert keeps its function because they only interact through that unified attention interface, which is a huge plus for flexibility.

Tom: Right, so we've got this architecture where you get these performance gains while maintaining the ability to swap out components easily. This paper, "SimWAM: A Simple World Action Model for End-to-End Autonomous Driving," shows how they achieve this without needing auxiliary motion modules.

Jane: It’s impressive how they manage to transfer those traffic dynamics priors directly into the action expert using just a joint flow matching approach, which is much simpler than building separate prediction modules and then connecting them later.

Lu: They show that the video expert can be initialized from models like Wan2 point 2-5B and map frames into latent tokens conditioned on the current frame, while future frames are "noised and reconstructed with flow matching," feeding that into the action expert as a trajectory velocity field prediction via flow matching.

Meng: That reliance on a pretrained video expert is key; it means the action expert doesn't have to learn everything from scratch regarding scene understanding; it inherits that knowledge. How does this affect the training stability?

Lalam: It seems like this approach could improve our culture of development because we can leverage massive amounts of pre-trained visual knowledge and apply it to driving tasks much more efficiently than building everything from zero.

Tom: Moving on to the improvements they suggest, SimWAM isn't just about achieving a result; it's about making the process smarter. They show that this setup allows for significant scaling because the video backbone can be swapped out without touching the action expert or changing the learning objective at all.

Title and authors: Jane: That scalability is something I find really appealing because it means if a newer, better video model comes out, we can plug it in and get an immediate performance boost for our planning system without rewriting the core logic.

Lu: Furthermore, they demonstrate that you can scale the action expert independently; for instance, increasing the Action DiT from 0 point 21B to 1 point 02B still improved PDMS from eighty-nine point nine to ninety point three without altering the video expert or the training objective at all, showing a very decoupled design principle here.

Meng: That parameter decoupling is what makes it attractive for deployment; we can tailor the action expert's size to meet specific latency targets while keeping the video representation powerful, which is exactly what we need in a production environment.

Lalam: It really emphasizes that AI systems don't always need monolithic designs; they can be composed of specialized parts that interact cleanly, which makes them much more adaptable to different needs.

Tom: And the final result they show on the NAVSIM benchmark is quite telling—achieving a PDMS of ninety-one point five with substantially lower latency compared to other world-model-based planners like DriveVLA-W0 (Flow-Matching).

Jane: That performance metric, ninety-one point five PDMS, when paired with that reduced inference latency, really paints a picture of a system that’s both highly capable and practical for real driving scenarios.

Lu: The paper also shows results from the isolated attention mask analysis where the final configuration achieved a PDMS of ninety point three in Table six which is solid validation across their experimental setups.

Meng: I want to touch on a limitation they mention; they state that while this method avoids explicit future-frame generation, it's still fundamentally relying on the training-time supervision signal from those video predictions to shape the observation representation used for planning.

Lalam: So, while inference is fast, the quality of that representation is inherently tied to how well that video prediction task was trained during the initial phase, which is something we need to keep in mind when deploying these systems.

Tom: That’s a fair point; it’s not magic that it bypasses future synthesis; it's just very smart training design and architecture that allows the learned priors to be used effectively at inference time.

Jane: So, if we look at the big picture, SimWAM suggests a pathway toward world models where the complex dynamics of driving are captured in a way that doesn't require heavy computation during operation.

Title and authors: Lu: It opens up possibilities for multimodal reasoning in driving where semantic knowledge from vision can inform action generation directly through this learned prior transfer mechanism.

Meng: For practical implementation, the modularity is the real selling point; having two separate experts that share an interface means we can iterate on the video part and the action part separately to optimize for different constraints.

Lalam: This work reinforces that for complex AI tasks, designing systems with clear separation of concerns, even if they interact through a unified interface, leads to more robust and adaptable solutions over time.

Tom: Well, Jane, Lu, Meng, Lalam—we’ve covered a lot about how SimWAM achieves high performance while cutting down on the computational burden of future-frame generation in autonomous driving.

Jane: It’s been really insightful hearing how they managed to decouple the training signal from the real-time inference process using that clever flow matching technique.

Lu: I think what’s most exciting is how they manage to leverage video priors without needing those extra, costly motion modules that used to be necessary in driving WAMs.

Meng: From an engineering standpoint, the ability to scale the action expert independently based on its required capacity while keeping the video backbone consistent is a huge factor for us looking at deployment budgets.

Lalam: It shows that AI development can be very modular, allowing us to build systems that are both powerful and adaptable as we learn more about what works best in practice.

Tom: So, let’s wrap up this discussion on "SimWAM: A Simple World Action Model for End-to-End Autonomous Driving." This paper provides a clear demonstration of how training signals can inform real-time action prediction efficiently through smart architectural choices like the isolated attention mask.

Jane: It’s a solid contribution to the world model space because it tackles the latency issue head-on by making the inference pipeline much leaner.

Lu: The implications for future research, I think, is that we can expect more models that focus on training supervision signals rather than relying solely on explicit future synthesis during deployment.

Meng: We need to keep watching how they handle those scenarios where the learned dynamics might not perfectly cover a novel situation, as that’s where any model will eventually show its limits.

Lalam: Overall, this paper is a great example of using sophisticated learning techniques to build systems that are inherently more efficient and better suited for the real-world challenges of autonomous driving.

The paper's summary: Tom: So, we’ve got the core of SimWAM down now—it’s about using training signals from future video prediction to teach an action expert directly, bypassing those slow imagination steps during driving.

Jane: Exactly; they're essentially saying the system can learn how to drive by observing what *would* happen in the future during training, which then makes real-time planning much faster because it doesn't have to generate those visuals on the fly.

Lu: I find that concept of co-training a video expert and an action expert through joint flow matching really fascinating from a theoretical standpoint; it creates this tight feedback loop where scene understanding directly shapes the learned motion.

Meng: From my side, what I’m focusing on is how they manage that isolation during inference; if the action prediction still needs to look ahead even once in real-time, we’re still facing latency issues regardless of the training setup.

Lalam: Looking at this from a cultural angle, it shows AI systems evolving past just reacting to the present and starting to internalize complex temporal dynamics, which fundamentally shifts how we think about autonomous behavior in our society.

Tom: That's right; they’ve engineered a way for the system to build that internal understanding during training so that when it hits the road, it can make decisions incredibly fast without needing those heavy future-frame synthesis modules.

Jane: And their results show they achieved a PDMS of ninety-one point five on NAVSIM, which is quite strong considering how much simpler their inference pipeline is compared to other world models we've seen recently.

Lu: The flexibility they built in, where the video backbone can be swapped out without breaking the action expert or changing the whole objective function, opens up huge possibilities for future research in multimodal AI.

Meng: That modularity is what keeps me interested; it means we can test different video representations against a fixed action expert and see how much performance we get for a given computational budget.

Lalam: It really reinforces that the most impactful advancements might come from designing these systems with clear, decoupled components that interact in specific ways rather than trying to build one giant, monolithic predictor.

Tom: Indeed, it seems SimWAM is proving that by smartly using training data to shape the representation for planning, we can get state-of-the-art results without incurring the heavy computational tax of explicit future image generation at runtime.

Jane: It’s a really neat trick—transferring those complex traffic dynamics priors from video generation into a lightweight action expert using joint flow matching is quite clever.

Lu: It suggests a path where vision and action learning are more deeply intertwined, moving beyond simple end-to-end connection toward truly shared representations between different AI modules.

Meng: If we can reliably scale the action expert independently while keeping the video representation high quality, that’s a huge win for deployment efficiency on edge devices.

Lalam: This development has implications for how we build trust in autonomous systems because it suggests a more intuitive and less computationally demanding way for them to make complex driving decisions.

Tom: So, while they’ve shown great metrics, the real excitement is in how they’ve engineered the inference path to be lean and efficient enough for actual deployment right now.

Jane: And that efficiency isn't just a number; it translates directly into a system that could be deployed much sooner in real-world testing scenarios where latency is critical.

Lu: We need to keep watching how this learned prior holds up when the driving environment gets unexpectedly chaotic, because relying on learned priors means we have to trust those priors implicitly in novel situations.

Meng: That's a valid concern; the paper mentions that while it performs well, the quality of that representation is still dependent on how well it was trained on specific traffic patterns.

Lalam: Ultimately, this work is a testament to how sophisticated training strategies can result in practical systems that are both powerful and adaptable for real-world applications.

The paper's improvements: Tom: So, we’ve been talking about how SimWAM gets its performance boost through joint flow matching, and now we're looking at what they suggest as improvements to push this system further into the real world.

Jane: It sounds like they’re not just happy with the results; they have a roadmap for making this model even more capable and practical for actual driving tasks.

Lu: I think one of their key suggestions is leveraging that decoupled design even more aggressively, allowing us to tailor the action expert's size independently based on how much planning capacity we need versus how much latency we can tolerate.

Meng: From an engineering standpoint, that’s exactly what I was thinking; being able to scale the action component separately means we can optimize for speed on my hardware while keeping a high-fidelity video representation, which is a huge practical advantage.

Lalam: It shows that the design itself is flexible enough to adapt to different constraints, which speaks volumes about how AI architectures should be built—they need to be modular and tunable for various deployment scenarios.

Tom: Right, it’s not just about getting one good number; they are proposing a design philosophy where you can swap out components for better performance or lower latency without having to rebuild the entire system from scratch.

Jane: That sounds like a very smart way to approach development because it lets us iterate on different parts of the system—the video expert and the action expert—independently while keeping the overall training objective consistent.

Lu: They are also pushing toward using Reinforcement Learning not just for imitation, but to directly optimize trajectory generation against complex driving quality rewards, moving beyond just following examples.

Meng: That’s where things get interesting; if RL can directly optimize for high-level driving quality in hard scenarios, we could see a significant improvement in how these systems handle those tricky situations that imitators often struggle with.

Lalam: This direction implies a future where AI systems don't just mimic human behavior but actively refine their driving logic through direct experience and optimization against real-world performance metrics.

Tom: So, they are moving from a model that just learns to *imitate* good driving to one that actively *optimizes* for high-quality driving maneuvers using reinforcement learning after the initial training phase.

Jane: That shift from imitation to direct optimization sounds like it could really help these systems become more robust when things get messy on the road.

Lu: The paper also hints at future work regarding how this learned prior handles novel scenarios, suggesting that research needs to focus on improving those learned representations for situations completely outside the training distribution.

Meng: That’s a realistic limitation they admit; relying on learned priors means we have to figure out how well that prior generalizes when it encounters something truly unexpected in the physical world.

Lalam: It underscores a vital cultural shift: building AI that can handle uncertainty and adapt its planning based on nuanced experience, rather than just relying on memorized scenarios.

Tom: So, the big idea is that SimWAM isn't just a finished product; it’s a foundation for future systems that can be optimized for real-world performance through direct learning methods.

Jane: It gives us hope that we can develop autonomous systems that are not just fast and accurate in known scenarios but also capable of handling the unpredictable nature of actual driving.

Lu: And with this level of modularity, I think we could see entirely new ways these vision-action models are combined for other complex tasks beyond just driving, opening up huge creative avenues.

Conclusion: Tom: So, to wrap up our discussion on "SimWAM: A Simple World Action Model for End-to-End Autonomous Driving," we’re looking at a model that smartly uses training signals from future video prediction to teach an action expert directly, bypassing those slow imagination steps during driving.

Jane: It really boils down to using the learned understanding of future scenes to make real-time decisions without needing heavy on-the-fly visual synthesis, which makes the system much leaner and faster.

Lu: The implications are huge for how we think about end-to-end systems; it suggests a path where vision and action learning are deeply intertwined through shared representations rather than being separate modules that just talk to each other.

Meng: From an engineering view, this means we can potentially deploy much more capable planning systems on less powerful hardware because the inference pipeline is streamlined, which is a massive win for practical deployment.

Lalam: For our culture as AI developers, it shows that we should focus on designing systems where the training process inherently builds in temporal understanding, which helps foster a more intuitive and robust way for AI to interact with complex environments.

Tom: Exactly; SimWAM’s success proves that smart training design can yield systems that are both highly capable and computationally efficient for real-world autonomous driving.

Jane: It gives us a lot of confidence in developing next-generation planning tools because we have a solid framework for transferring those rich traffic dynamics priors into action experts efficiently.

Lu: I think the modularity they introduced is particularly exciting, suggesting that future research could easily combine this structure with other modalities like audio or sensor data without needing to completely rewrite the core learning objective.

Meng: That ability to swap components independently really means we can rapidly prototype and test different architectural ideas for action experts without having to redo all the vision processing pipeline.

Lalam: It highlights that AI systems should be built with a mindset of adaptability, allowing us to easily plug in new visual backbones or scale up the action component as needed based on our specific needs.

Tom: So, we’ve seen how SimWAM leverages joint flow matching and isolated attention masks to achieve strong performance while significantly reducing the computational overhead of future-frame generation during inference.

Jane: It's a really neat way to keep the system powerful without making it too slow for practical use on roads.

Lu: This paper opens up possibilities for integrating world models with other complex reasoning tasks where temporal context and visual scene understanding are crucial, which could lead to some very creative applications down the line.

Meng: I just want to make sure we keep an eye on those scenarios where the learned dynamics might not perfectly cover a totally new type of driving situation, because that’s where any model will eventually show its limits.

Lalam: That uncertainty is precisely what drives our next phase of development; it pushes us to build systems that are not just good at known tasks but also resilient and capable of adapting to the unknown.

Tom: Well, that’s a great summary of SimWAM: A Simple World Action Model for End-to-End Autonomous Driving. Thanks to Lu, Meng, and Lalam for those excellent insights!

Jane: It’s been wonderful discussing this paper with you all; it really shows how clever training methods can lead to practical improvements in autonomous driving performance.

Lu: I look forward to seeing how the community builds on this modular approach in the coming years.

Zongchuang Zhao, Xin Zhou, Tianyang Xu, Zhengyang Sun, Kaixuan Zhou, Yu Wu, Honglin Li

Huazhong University of Science & Technology

cs.CV

Submitted: 2026-08-07

Updated: 2026-09-29

Comments: The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/

Code: https://github.com/H-EmbodVis/SimWAM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: SimWAM presents a novel World-Action Model (WAM) designed to improve end-to-end autonomous driving by leveraging future video prediction as a training signal, thereby avoiding costly test-time future

Key concepts

World Action Model (WAM)
A model designed for end-to-end autonomous driving that improves performance by leveraging future video prediction as a training signal, aiming to avoid the need for costly test-time future imagination during actual driving.
Joint Flow Matching Objective
A method used to separate video prediction and action prediction. It involves defining a combined training objective, L = L act_FM + lambda L vid_FM, which balances the signals from both experts during learning.
Isolated Attention Mask
A structural feature in SimWAM that keeps the action prediction independent of future frames during real-time operation. This allows for direct trajectory prediction from current inputs without needing explicit future-frame generation.

Terminology

Summary

SimWAM presents a novel World-Action Model (WAM) designed to improve end-to-end autonomous driving by leveraging future video prediction as a training signal, thereby avoiding costly test-time future imagination. This method co-trains a pretrained video expert and a lightweight action expert using joint flow matching. By employing an isolated attention mask, SimWAM achieves direct trajectory prediction at inference without needing explicit future-frame generation, offering substantial latency reduction while transferring traffic dynamics priors effectively.

Model Architecture and Training Objective

SimWAM is structured around two distinct experts: a pretrained video expert and a lightweight action expert. The video expert is initialized from models like Wan2.2-5B and utilizes a VAE to map driving frames into latent tokens, conditioned on the current frame, with future frames being noised and reconstructed with flow matching. The action expert is a lightweight Diffusion Transformer that predicts the trajectory velocity field via flow matching, conditioned on the current observation latents. The joint training objective is defined as:

L = L act FM + λL vid FM,

where both terms instantiate Eq. 1 (flow matching) over the action trajectory and future-frame latents, with λ balancing the two components. This co-training allows future-scene prediction to shape the observation representation used for planning.

Inference Mechanism and Latency Reduction

A key innovation of SimWAM is its inference pipeline, which avoids explicit future-frame generation. The action expert is conditioned on a representation produced from the current observation, denoted as z(ot). This allows for a simple policy interface:

pθ(at+1:t+H ot, st, l) = pθ(at+1:t+H z(ot), st, l)

This decoupling is achieved through an isolated attention mask that keeps action prediction independent of future frames. This structural modification means the future-video prediction objective serves as a training-time supervision signal, and at inference, the model directly predicts trajectories from the current inputs, thus bypassing explicit future-frame prediction and substantially reducing inference latency compared to methods that follow an imagine-then-act pipeline.

Optimization via Reinforcement Learning

The imitation learning stage is followed by reinforcement learning (RL) to optimize trajectory generation directly toward driving quality. Since the deterministic flow ODE lacks the stochasticity needed for exploration, SimWAM reformulates it as a stochastic SDE (Eq. 2). The action expert is then reinforced using Flow-GRPO, which involves sampling a group of G candidate trajectories and evaluating them using the compositional NAVSIM PDM reward. This process allows for diverse maneuver exploration and direct optimization of a compositional driving reward, focusing RL on hard navtrain scenarios with the lowest PDMS after imitation learning.

Flexibility and Scalability

SimWAM exhibits significant flexibility in both architecture and model scale due to its decoupled design. The two experts share no parameters and interact only through a unified attention interface, allowing the video backbone to be replaced without modifying the action expert or inference pipeline. Furthermore, the action expert can be scaled independently; for instance, increasing the Action DiT from 0.21B to 1.02B steadily improves PDMS from 89.9 to 90.3 without changing the video expert or training objective. This parameter-decoupled two-expert design supports different performance and computation budgets through one unified design, accommodating stronger video priors while allowing adaptation of the action expert for efficiency requirements.

Performance and Generalization

Experiments on the NAVSIM benchmark demonstrate SimWAM's effectiveness, achieving a PDMS of 91.5 with substantially lower latency than state-of-the-art world model planners. The method also achieves zero-shot to nuScenes performance without fine-tuning, indicating architectural scalability and cross-domain generalization. Qualitative results show that after reinforcement, the model follows the intended route more decisively and completes a larger portion of each maneuver, confirming that internalized dynamics priors benefit planning beyond explicit future synthesis at inference. The final configuration achieves the best PDMS of 90.3 in Table 6 for the isolated attention mask analysis.

Key Findings Summary

  1. SimWAM effectively transfers traffic dynamics priors from a pretrained video generator to the planner without auxiliary motion modules, using joint flow matching.

  2. The isolated attention mask ensures the action expert remains independent of future-frame representations, enabling efficient inference without explicit future-frame generation.

  3. RL optimizes the action expert for direct continuous trajectory prediction after co-training, improving PDMS to 91.5 by directly optimizing driving quality beyond trajectory imitation.

  4. The architecture is flexible: the video backbone can be replaced, and the action expert can be scaled independently based on desired planning capacity and latency targets.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed SimWAM: A Simple World Action Model for End-to-End Autonomous Driving. The core innovation lies in decoupling future-frame generation from action prediction during inference by leveraging training-time supervision via flow matching, while maintaining the benefits of world models.

Here are the specific improvements that can be made to existing AI systems and what those improved systems can achieve:


)

  1. A new end-to-end autonomous driving planner that achieves state-of-the-art performance (91.5 PDMS on NAVSIM) with significantly lower inference latency than current world model planners.

  2. The ability to transfer sophisticated traffic dynamics priors learned from large, pretrained video generation models (like Wan2.2-5B or Cosmos) directly into a lightweight action expert without requiring any auxiliary motion modules or explicit future-frame synthesis during real-time operation.

  3. A system capable of zero-shot generalization to new driving benchmarks (e.g., nuScenes) without any fine-tuning, relying solely on the learned dynamics prior transferred from the training set.

  4. A scalable architecture where the video expert backbone can be swapped out for newer or more domain-relevant video generation models (e.g., replacing LTX-Video with Cosmos2.5) without needing to modify or retrain the action expert or the core learning objective.

  5. A system that optimizes trajectory generation directly toward high-level driving quality (beyond mere expert imitation) by incorporating Reinforcement Learning, enabling the exploration of diverse, safe maneuvers in complex, low-PDMS scenarios (navhard).

  6. A planning pipeline that maintains high safety metrics—specifically superior No Collision (NC) and Time-to-Collision (TTC)—even when operating under reactive conditions where surrounding agents respond to the ego vehicle's trajectory.

Abstract

In autonomous driving, World-Action Models (WAMs) have improved end-to-end planning by transferring video dynamics priors to action prediction, but many still couple planning with future-video generation at inference, incurring substantial computational overhead. We present SimWAM, a simple yet effective WAM that leverages future-video prediction solely as a training-time supervision signal. It co-trains a pretrained video expert and a lightweight action expert with joint flow matching. An isolated attention mask keeps action prediction independent of future frames, allowing trajectory prediction without future-frame generation at inference. This design supports multiple pretrained video backbones and independent action-expert scaling within a shared attention interface, while preserving the joint learning objective. Moreover, we apply reinforcement learning to optimize a compositional driving reward beyond trajectory imitation. Experiments show that SimWAM achieves 91.9 PDMS on NAVSIM with a favorable trade-off between accuracy and latency among world-model-based planners, while transferring zero-shot to nuScenes. It also achieves competitive planning accuracy on WOD-E2E and PhysicalAI-Autonomous-Vehicles. These results position SimWAM as a plain yet solid baseline for efficient autonomous driving. The code and model weights are available at https://github.com/H-EmbodVis/SimWAM/.

Sources

Related papers