GeoWAM: Visual Geometry World Action Models for Autonomous Driving

arXiv:2608.23486 · cs.CV, cs.RO · Submitted 2026-08-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "GeoWAM: Visual Geometry World Action Models for Autonomous Driving".

Jane: GeoWAM introduces a novel visual geometry world action model designed for autonomous driving by shifting the focus from predicting future images to forecasting future scene geometry.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at the paper "GeoWAM: Visual Geometry World Action Models for Autonomous Driving," which sounds really technical, but it’s about how we model driving. The core idea is shifting the focus from predicting what the next picture will look like to forecasting how the physical space itself will change over time.

Jane: That makes sense when you think about driving; you're not just predicting a new frame, you're trying to understand where things are going in three dimensions, which GeoWAM tries to do by using point clouds instead of just pixels.

Lu: The authors are Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino, Dhruva Dixith Kurra, Min Cai, Tom Lampo and Yu Yin. They bring a really solid group of expertise to this kind of spatial reasoning.

Meng: I'm curious about the immediate impact of focusing on geometry versus pixels; from an engineering standpoint, if we can get a state representation that’s inherently three dee, it simplifies the downstream planning task significantly.

Lalam: I think what excites me most is how this moves us closer to building truly intuitive AI systems because geometry aligns so well with the physical world where actions happen, which could really improve how we design those driving policies.

The paper's summary: Tom: Well, the paper explains that existing World Action Models often rely on video generation backbones to predict future observations and then an action head to guess the ego-trajectory. GeoWAM tackles this by arguing that pixels are just an indirect way of seeing dynamics because they mix geometry with things like lighting and texture.

Jane: So, what GeoWAM does is build a model that forecasts the actual future three dee scene structure based on historical multiview observations, which then provides a much clearer foundation for predicting the ego vehicle's motion.

Lu: They introduce this visual geometry world model pretraining stage where they encode image sequences into historical tokens and use a future geometry decoder to predict the three dee structure, which explicitly shows the spatial transformations happening in the scene.

Meng: I see them using components like a learned query seed for every future step and location, along with causal temporal self-attention and cross-attention mechanisms to connect this predicted geometry back to historical data. That sounds computationally intensive but necessary if it's truly capturing those complex spatial dynamics.

Lalam: It’s interesting that they use a shared geometry head that decodes this into a dense point map and a confidence map, which predicts the geometric evolution without needing to reconstruct the future image appearance, which feels like a big efficiency win for pure planning.

The paper's improvements: Tom: The real improvement they propose is using an "inverse-dynamics-like formulation" to extend this geometry world model into a full action model. This means they don't just predict the scene; they use that predicted geometry to infer the future ego motion, and then map that motion directly onto a trajectory.

Jane: That’s where the magic happens, because instead of guessing movement based on visual cues in an image, you are deriving the movement directly from the predicted three dee structure of where you're going.

Lu: They couple future ego-token decoding with trajectory decoding, where the action head takes those deepest historical and predicted ego tokens and refines them using a causal temporal transformer before a learned trajectory query maps it to a single future trajectory.

Meng: From an engineering standpoint, this structure seems robust because it separates the scene forecasting from the motion prediction, which might make debugging much clearer when things go wrong in complex scenarios.

Lalam: If we can get that spatial alignment right, we’re talking about a system that handles complex maneuvers like turning much better than current image-based methods because the physics are baked into the representation.

Conclusion: Tom: So, to wrap up on "GeoWAM: Visual Geometry World Action Models for Autonomous Driving," the paper shows how shifting from pixel prediction to geometry forecasting provides a state space that is much more natural for driving dynamics, leading to better trajectory predictions through that inverse-dynamics approach.

Jane: Essentially, they’ve shown how pretraining a visual geometry world model allows us to capture the spatial structure and temporal dynamics of scenes in a way that directly informs ego motion planning.

Lu: The implication is that we can achieve more physically grounded driving policies because the model is explicitly aware of the three dee space where actions are executed, which is something previous video-based models struggled with.

Meng: From a practical standpoint, this means our perception pipeline could feed a cleaner geometric understanding directly into the planner, potentially reducing reliance on messy visual features for trajectory generation.

Lalam: I think this work really validates the idea that geometry isn't just extra data; it’s a fundamental state representation for autonomous driving AI. We should see systems using this approach become much more reliable in real-world driving situations.

Yiren Lu, Xin Ye, Jiaming Liu, Philip Jacobson, Jin Yao, Yi-chung Chen, Liam Merino

Uber AV Labs

cs.CV, cs.RO

Submitted: 2026-08-24

Updated: 2026-09-29

Code: https://github.com/OpenDriveLab/OpenScene

Importance score: 90/100

The gist: GeoWAM introduces a novel visual geometry world action model designed for autonomous driving by shifting the focus from predicting future images to forecasting future scene geometry.

Key concepts

Visual Geometry World Model
This is the first stage where the system learns to predict what a future scene will look like in 3D space, using past camera images as input. It uses geometry encoders to create tokens that describe the spatial structure of the environment, allowing it to forecast future point clouds without needing to generate new pictures.
Ego Token Decoding
This process takes learned seeds representing the vehicle's current state and combines them with predicted future geometry information. It creates 'ego tokens' that explicitly describe how the vehicle should move in the next few steps, ensuring the predicted motion matches the forecasted physical changes in the scene.
Inverse-Dynamics-like Formulation
This is a method used to connect scene geometry prediction with driving action planning. Instead of predicting actions from scratch, it uses a formulation that infers future ego motion by observing how the predicted 3D geometry evolves over time, directly linking the physical environment's changes to the required vehicle trajectory.
Point Cloud State Space
Instead of using pixels (which are indirect), this approach uses explicit 3D point clouds as the state representation. This is significant because it naturally captures the rigid and non-rigid spatial transformations governing a scene, which directly corresponds to how a physical vehicle moves in space.

Terminology

Summary

GeoWAM introduces a novel visual geometry world action model designed for autonomous driving by shifting the focus from predicting future images to forecasting future scene geometry. This approach is significant because it leverages explicit spatial structure, represented by point clouds, as a more natural state space for modeling scene evolution and aligning directly with the physical space where driving actions are executed. By pretraining on historical multiview observations to predict future 3D structure, GeoWAM provides a geometrically grounded foundation that explicitly captures the spatial transformations governing scene dynamics.

Motivation and Problem Statement

Existing World Action Models (WAMs) often learn scene dynamics implicitly by combining video-generation backbones for future-observation prediction with an action head for ego-trajectory prediction. The paper argues that pixels provide only an indirect representation of these dynamics, entangling geometry and motion with appearance, texture, and illumination, which forces the model to infer three-dimensional transformations from two-dimensional observations. In contrast, geometry explicitly captures spatial structure and the rigid and non-rigid transformations that govern scene evolution while directly aligning with the space in which driving actions are executed. Therefore, pixels are not an ideal state representation for modeling driving dynamics.

Methodology: Visual Geometry World Model Pretraining

The training of GeoWAM consists of two stages. The first stage involves pretraining a visual geometry world model to predict future scene geometry from historical multiview images. This process utilizes a geometry encoder (like DVGT-2) to encode the multiview image sequence into multi-level historical tokens, which are then processed by a future geometry decoder. Key components of this decoding include:

  1. A learned query seed for every future step, camera view, and spatial location to construct future geometry queries.

  2. A causal temporal self-attention mechanism to model the evolution of each spatial location across the F future steps.

  3. Cross-attention mechanisms that connect the updated queries to historical memory and predicted future geometry tokens to retrieve relevant context.

  4. A shared geometry head (Gψ) that decodes these representations into a dense point map and a per-pixel confidence map, predicting geometric scene evolution without reconstructing future image appearance.

Methodology: GeoWAM World Action Model Extension

After geometry pretraining, the model is extended to trajectory planning via an inverse-dynamics-like formulation. This stage couples the learned geometric dynamics with action prediction. The process involves:

  1. Future ego-token decoding, where learned ego-query seeds are used to construct queries that cross-attend to both historical geometry memory and predicted future geometry tokens to produce future ego tokens that describe ego motion consistent with the forecast scene evolution.

  2. Trajectory decoding, where the action head takes the deepest-level historical ego tokens and predicted future ego tokens, appends them along the temporal dimension, and refines them with a causal temporal transformer. A learned trajectory query then cross-attends to these refined features to map directly to a single future trajectory.

Training Objectives and Evaluation

The complete finetuning objective for planning is defined as:

(Lplan) = Lpre + λtrajLtraj + λposeLpose.

Where:

  1. Future geometry pretraining objective (Lpre) combines feature-level alignment (cosine distance between predicted features and targets) and point-map supervision, specifically combining Euclidean point regression, confidence-aware regression, and multi-scale surface-normal consistency.

  2. Trajectory supervision uses an L1 regression loss (Ltraj).

  3. An auxiliary L1 loss (Lpose) predicts the relative poses between historical frames.

GeoWAM is evaluated across multiple settings: future-geometry prediction on the nuScenes validation set and open- and closed-loop planning on NAVSIM, using metrics such as absolute relative error, threshold accuracy δ < 1.25 for geometry prediction, and the Extended Predictive Driver Model Score (EPDMS) for planning. Open-loop evaluations show that GeoWAM yields substantially stronger driving policies than image-based alternatives, while closed-loop planning demonstrates its effectiveness in handling accumulated deviations.

Key Contributions

The primary contributions of GeoWAM are:

  1. Motivating geometry as a native state representation for world action models, directly aligning scene dynamics with the three-dimensional space in which driving actions are defined.

  2. Pretraining a visual geometry world model to forecast future scene geometry from historical multiview observations, enabling it to capture the spatial structure and temporal dynamics of driving scenes.

  3. Introducing GeoWAM, which extends this pretrained model into a world action model through an inverse-dynamics-like formulation that infers future ego motion from predicted geometry and maps it to an ego trajectory.

  4. Validating GeoWAM through future-geometry prediction and open- and closed-loop planning, demonstrating the effectiveness of visual geometry world modeling for autonomous driving.

Improvements for AI systems

Here are the specific improvements that can be made to existing AI systems, based on the GeoWAM paper, and what those improved systems can achieve:


The core improvement lies in shifting the state representation from pixel space (images) to geometry space (point clouds), which is more naturally aligned with physical driving dynamics.

Here are the specific improvements and capabilities:

  1. The system can be pretrained using historical multiview images to learn a model that forecasts future 3D scene geometry directly, rather than just future RGB pixels.

  2. This learned geometric dynamics can then be used to predict the ego vehicle's future trajectory by leveraging an inverse-dynamics-like formulation (GeoWAM).

  3. The resulting system can perform high-fidelity, long-horizon future geometry prediction, accurately forecasting the 3D structure of the environment several seconds into the future.

  4. Because planning is conditioned on this explicit 3D spatial structure, the ego vehicle's action predictions become inherently more physically grounded and spatially coherent.

  5. The improved system can perform robust open-loop and closed-loop trajectory planning by directly using predicted geometry to inform motion planning, leading to superior performance in complex maneuvers like turning and lane keeping.

This improved AI system (GeoWAM) can achieve the following specific capabilities:

  1. It will exhibit substantially stronger driving policies compared to image-based alternatives because it explicitly models spatial structure and rigid/non-rigid transformations that govern scene evolution.

  2. It can perform precise metric-scale reconstruction of future scene geometry (e.g., predicting exact 3D point maps) rather than just visually plausible future images.

  3. It will provide a spatially grounded representation for world action modeling, meaning the learned dynamics are directly aligned with the 3D space in which driving actions are defined, improving safety and control during maneuvers where spatial context is critical (e.g., turning left/right).

  4. It can effectively handle complex planning tasks by integrating both future scene forecasting and ego-trajectory generation within a shared geometric space, resulting in higher EPDMS scores on challenging benchmarks like NAVSIM's two-stage planning protocol.

Sources

Related papers