Grounded World Model: Latent Planning with Language Goals

arXiv:2604.11751 · cs.RO, cs.AI · Submitted 2026-04-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Grounded World Model: Latent Planning with Language Goals".

Dev: This work proposes a Grounded World Model (GWM) to enhance planning in Model Predictive Control (MPC) by leveraging a vision-language-aligned latent space,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So, let's talk about the title and authors of this paper, Grounded World Model: Latent Planning with Language Goals. The title itself really tells you what’s happening—they are grounding their world model using language goals to guide planning within a latent space.

Dev: I agree, Rosa; it highlights that they aren't just looking at the robot's immediate surroundings but are tying the action selection directly to a high-level linguistic instruction, which is a significant shift from traditional methods.

Taro: The authors include researchers from EPFL and UofT, which suggests they’re bringing together expertise in both robotics and large-scale foundation models, which is a big deal for this kind of work.

Rosa: Exactly; it shows the multidisciplinary nature required to bridge the gap between high-level language understanding and low-level motor control planning effectively. It points toward a more integrated approach to building autonomous systems.

Dev: And considering their focus on using a vision-language-aligned latent space, it suggests they are leveraging existing powerful models rather than trying to build everything from scratch, which makes sense for rapid progress in this area.

Taro: I'm interested in how this title sets the stage for the research; it implies that the core innovation isn't just a new model architecture, but rather a novel way to *apply* existing vision-language knowledge to planning.

Rosa: Right, Taro; so it’s not just about a new algorithm, but about reimagining how we use these powerful foundation models for complex visuomotor tasks. It’s about using language as the primary steering mechanism for physical movement.

Dev: I think the implication is that future planning systems won't just be optimizing trajectories based on visual features, but will be optimizing them based on a semantic understanding of what the human actually wants to achieve.

Taro: That semantic steering capability could lead to agents that are much more capable of handling instructions that are phrased in nuanced or abstract ways, not just simple geometric commands.

Rosa: So it’s about moving planning from purely visual matching to a language-guided outcome prediction, which opens up a whole new way for agents to interpret goals.

Dev: It really changes the problem space from finding the closest image to finding the most semantically aligned future state, which is a much richer target for optimization.

Taro: That transition suggests that if we can solve this well, we might see agents demonstrating capabilities in task completion that were previously considered too abstract or complex for standard reinforcement learning setups.

Rosa: It’s about making the world model itself an intelligent planner guided by language, rather than just a reactive predictor of visual states.

Dev: And that makes sense because it shifts the computational burden from generating perfect goal images to calculating semantic similarity in a latent space, which sounds like a trade-off worth making for better generalization.

The paper's summary: Rosa: Moving on to the actual summary of Grounded World Model: Latent Planning with Language Goals, the authors explain that their core method is to learn this GWM within a vision-language-aligned latent space. They are training a model like Qwen3-VL-Embedding to encode both images and text into one shared embedding space.

Dev: So, the summary explains that this latent space lets them compute cosine similarity between the predicted future outcome embeddings of an action proposal and an instruction embedding derived from that same foundation model. That’s the mechanism driving their planning strategy.

Taro: The summary mentions they learn a transition function within this latent space without changing the weights of the foundation model, which is important because it means they are preserving a lot of the existing world knowledge from the pretraining phase.

Rosa: That preservation of world knowledge is key; it implies that each proposed action gets scored based on how close its future outcome is to what a natural language instruction actually describes in that shared space.

Dev: So, instead of training an entirely new world model from scratch, they are fine-tuning the transition function in this latent space, which should make the learning process much more stable and less prone to catastrophic forgetting.

Taro: The summary also mentions that during inference, actions are executed sequentially until the task instruction is completed in a closed-loop rollout until completion.

Rosa: That closed-loop rollout confirms that the system isn't just making a single guess; it’s actually executing the plan step by step, checking its progress against the instruction continuously.

Dev: This sequential execution model gives me some thoughts on latency; if they are generating future embeddings sequentially rather than in parallel, that could explain why inference efficiency isn't as fast as some VLA baselines.

Taro: I wonder if that sequential nature is actually a feature here, perhaps ensuring better fidelity for the prediction of each subsequent state given the previous one.

Rosa: It seems like they are balancing the need for deep semantic understanding with the operational reality of sequential execution, which is something we have to consider when we think about deployment in a robot setting.

Dev: So, in short, they’re using latent space similarity to score actions against language instructions to guide a sequence of actions that must be executed until the task specified by the natural language is finished.

Taro: That makes sense; it’s not just about prediction anymore; it’s about generating an action sequence that satisfies a specific linguistic condition throughout its execution.

The paper's improvements: Rosa: Now we get into the specific improvements the authors highlight, and they point out that their approach surpasses existing VLM-based VLAs in semantic generalization, showing an eighty-seven percent success rate on the WISER benchmark compared to a twenty-two percent for traditional VLAs.

Dev: An eighty-seven percent success rate is quite high, Rosa; that’s a substantial jump over the baseline of twenty-two percent, suggesting that their method is much more robust when dealing with unseen visual signals and referring expressions.

Taro: That performance gap on the WISER benchmark seems to be the most concrete evidence they have for this semantic generalization claim; it’s not just theoretical—it’s measured against a set of twenty-four distinct world knowledge categories.

Rosa: Indeed, Taro; the fact that these tasks require world knowledge and referring expressions unseen during training makes this result much more impressive because it proves the system isn't just memorizing specific visual correlations.

Dev: They also mention that they have prompt decomposition for GWM-MPC further boosting performance compared to VLM-based VLAs, which suggests that breaking down the instruction helps the model leverage the full language understanding capability.

Taro: If prompt decomposition is effective, does it mean we can expect this improved handling of complex instructions to translate into agents being able to handle more intricate, multi-part commands in practice?

Rosa: I think so; if they can decompose a complex instruction into smaller sub-tasks for scoring, the agent should be better equipped to sequence those sub-goals correctly.

Dev: From an engineering view, this improved handling of compositionality is crucial because it means the system isn't just relying on one giant language prompt, but a structured way of breaking down the planning problem.

Taro: That structure in instruction handling could be what allows for better safety margins when dealing with unexpected world states during execution.

Rosa: So they are showing that this method doesn't just generalize; it handles compositionality and referring expressions much more effectively than previous models did on those complex benchmarks.

Dev: But I still have that efficiency concern from earlier; if the system is sequential, we need to make sure the latency doesn't become a bottleneck when we move towards real-world interaction speeds.

Conclusion: Rosa: So, to wrap up with these points from Grounded World Model: Latent Planning with Language Goals, it seems the authors have successfully shown how grounding action selection in a vision-language latent space using language goals can lead to much higher success rates on semantic generalization tasks compared to older VLA approaches.

Dev: It really boils down to using the GWM architecture to transform MPC into a VLA system where the planning score is based on semantic closeness rather than just visual distance, which is a significant methodological change.

Taro: I think the big implication here is that we are moving towards agents that can genuinely interpret and execute complex, natural language goals in ways that feel much more intuitive to interact with.

Rosa: That sounds like the direction we’re heading—agents that don't just follow hard-coded paths but truly understand the intent behind what they are asked to do.

Dev: I think the system provides a solid framework for integrating language into robotic control loops in a way that prioritizes semantic alignment in decision-making, even if inference speed is currently slower than some alternatives.

Taro: Moving forward, we need to keep watching how they address those deployment questions regarding long-term stability and real-world robustness so this capability translates into something practical for the field.

Rosa: That’s a fair call; the Grounded World Model: Latent Planning with Language Goals offers a very promising direction for how we can make our robotic agents smarter planners.

Dev: Well, it's definitely worth keeping an eye on this work as they continue to refine the system for real-world deployment and performance tuning.

Taro: I agree; the potential for robust semantic generalization is certainly something that warrants a lot of attention from everyone in autonomy research.

EPFL

cs.RO, cs.AI

Submitted: 2026-04-13

Updated: 2026-09-30

Code: https://github.com/QuanyiLi/gwm-wiser

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 83/100

The gist: This work proposes a Grounded World Model (GWM) to enhance planning in Model Predictive Control (MPC) by leveraging a vision-language-aligned latent space, offering a novel approach to semantic

Key concepts

Grounded World Model (GWM)
A model that learns within a vision-language-aligned latent space. It encodes both images and text into a shared embedding space, allowing the system to compute similarity between predicted future outcomes and language instructions.
Latent Space Similarity
The mechanism used for planning where cosine similarity is computed between the embeddings of an action proposal's future outcome and an instruction embedding. This scores actions based on how close their results are to what the natural language goal describes.
Semantic Generalization
The ability of the model to perform well on unseen visual signals and referring expressions, measured against world knowledge categories. The paper shows a high success rate (87%) compared to traditional vision-language models (22%).
Prompt Decomposition
A technique used in GWM-MPC that breaks down complex instructions into smaller sub-tasks for scoring. This helps the agent handle multi-part commands better and improves compositionality.

Terminology

Summary

This work proposes a Grounded World Model (GWM) to enhance planning in Model Predictive Control (MPC) by leveraging a vision-language-aligned latent space, offering a novel approach to semantic generalization for visuomotor agents. The method addresses the challenge of obtaining goal images in advance by scoring proposed actions based on the similarity of their predicted future outcomes embeddings to natural language instructions, transforming MPC into a Vision Language Action (VLA) system that surpasses existing VLM-based VLAs in semantic generalization, as demonstrated on the WISER benchmark.

Core Concept: Grounded World Model (GWM)

The GWM is designed to operate within the latent space of a pretrained multi-modal retrieval model, specifically Qwen3-VL-Embedding. This foundation model encodes images and text into a shared embedding space where cosine similarity can be computed. The GWM learns the transition function in this latent space without altering the foundation model's weights, which largely preserves the multi-modal world knowledge of Qwen3-VL-Embedding. This allows each proposed action to be scored based on how close its future outcome is to the task instruction, reflected by the similarity of embeddings.

Planning Mechanism: WISER Benchmark and MPC

The system utilizes Model Predictive Control (MPC) where a batch of candidate trajectories is proposed, and the trajectory yielding the minimum cost is executed. Trajectories are proposed using a K-Nearest Neighbors (KNN) approach to retrieve demonstrated actions from a dataset D, which serves as the continuous proxy for planning. The scoring step involves calculating cosine similarity between the predicted future outcome embeddings and an instruction embedding derived from the same foundation model. Specifically, it selects the action sequence whose predicted future outcome embedding exhibits the highest cosine similarity with the instruction embedding.

Action Tokenization: Rendering-based Action Tokenization (RAT)

To encode both current observation and proposed actions for prediction, the paper employs Rendering-based Action Tokenization (RAT). This method involves sequentially rendering proposed actions into images using the third-person main camera parameters and the robot’s URDF. This approach allows the vision encoder of Qwen3-VL-Embedding to extract features without introducing additional learnable parameters, making it embodiment-agnostic and enabling zero-shot generalization to the xArm6 robot.

Semantic Generalization Evaluation: WISER Benchmark

To rigorously evaluate semantic generalization, the World-knowledge Integrated Semantic Embodied Reasoning (WISER) benchmark is introduced. This benchmark consists of 24 subsets corresponding to distinct categories of world knowledge (e.g., numbers, food, animals). Test tasks are constructed with world knowledge and referring expressions that are unseen during training, yet the motions required to complete them are already demonstrated during training. The success rate gap between traditional VLAs (average success rate of 22% on test) and GWM-MPC (87% success rate on test) highlights the superior semantic generalization of the proposed approach.

Key Advantages and Findings

The GWM-MPC system demonstrates strong performance, achieving an 87% success rate on the test set comprising 288 tasks that feature unseen visual signals and referring expressions. Furthermore, ablation studies confirm robustness:

  1. GWM largely preserves world knowledge because it leverages the pretrained latent space to learn the transition function without altering the foundation model.

  2. The system is robust to hyperparameters, with performance bottlenecked by the foundation model.

  3. RAT enables zero-shot generalization to a new embodiment with different action spaces, kinematics, and appearance (e.g., xArm6).

  4. Prompt decomposition for GWM-MPC further boosts performance compared to VLM-based VLAs, suggesting that GWM can leverage the intact language understanding ability of the foundation model.

Inference Efficiency

While VLA baselines generally exhibit better inference efficiency than GWM-MPC on test tasks, this is attributed to the fact that GWM requires forwarding the GWM N = 12 times to get future embeddings for all proposals, and because it generates future embeddings sequentially rather than in parallel due to Qwen encoder batching issues. The system's training efficiency is noted as being computationally efficient, requiring only 20 GPU hours on our proposed WISER benchmark.

Score Function Design Details

The scoring function utilizes sub-task embeddings derived from decomposing the instruction into two parts: a pick sub-task and a place sub-task. These are encoded using Qwen3-VL-Embedding with the initial observation and current observation to produce sub-task embeddings: z pick g and z place g. The final selection score is determined by comparing candidate future embeddings against these specific sub-task instructions, ensuring that the action sequence selected aligns with both the grasping and placement goals.

Performance Upper Bound

The performance of GWM is bounded by the Qwen3-VL-Embedding.

Improvements for AI systems

As a fastidious researcher, I have analyzed this paper, Grounded World Model for Semantically Generalizable Planning, and identified several high-impact areas for improvement in existing AI systems. The core contribution is the Grounded World Model (GWM) approach integrated into Model Predictive Control (MPC), leveraging a vision-language-aligned latent space.

Here are the specific improvements and capabilities this system enables:


)

Improvements to AI Systems Enabled by GWM-MPC:


  1. 】Semantic Generalization Beyond Training Set Overfitting (Zero-Shot Robustness):

  2. 】Goal Specification via Natural Language in Latent Space (Human-Friendly Interface):

  3. 】Action Selection via Semantic Similarity Scoring (Robust Policy Guidance):

  4. 】Cross-Embodiment Generalization without New Parameters (Robot Agnostic Deployment):

  5. 】Improved Handling of Compositional and Referring Expressions (Complex Instruction Following).

  6. Semantic Generalization Beyond Training Set Overfitting:

The system can achieve significantly higher success rates on unseen test tasks compared to traditional VLAs. By grounding predicted future outcomes in a shared vision-language latent space (via Qwen3-VL-Embedding), the agent learns to associate predicted trajectories with high semantic similarity against the natural language instruction embedding, rather than relying on overfitting to specific visual correlations or training data shortcuts.

  1. Goal Specification via Natural Language in Latent Space:

The system allows for goal specification using natural language instructions (e.g., Pick up the [X] and place it onto the [Y]) directly within the world model's latent space scoring mechanism, bypassing the need to generate a specific goal image beforehand. This makes task instruction an intrinsic part of the planning objective, which is more flexible and intuitive than relying solely on image-based goals.

  1. Action Selection via Semantic Similarity Scoring:

The MPC framework uses a sophisticated scoring function based on cosine similarity between predicted future embeddings and the instruction embedding (derived from Qwen3-VL-Embedding). This ensures that the robot selects actions whose predicted outcomes are semantically closest to the user's intent, leading to more semantically coherent and goal-directed behavior than simple distance metrics (like MSE) in latent space.

  1. Cross-Embodiment Generalization without New Parameters:

The system is designed with an embodiment-agnostic rendering-based action encoder (RAT). This allows the learned world model to be reused for zero-shot generalization to different robot morphologies, kinematics, and appearances (e.g., from a Panda robot to an xArm6 robot) simply by retraining the action encoder on the new embodiment's data. This avoids the need for full fine-tuning of the policy or world model backbone for every new robot platform.

  1. Improved Handling of Compositional and Referring Expressions:

By leveraging the foundation model's inherent visual-language understanding, GWM-MPC can robustly handle novel visual signals and referring expressions that were not seen during training, provided the underlying motions required to complete the task have been demonstrated previously. The decomposition of tasks into sub-tasks (as shown in Table 3) further proves its ability to handle compositional semantics better than VLAs that rely on fine-tuned sentence structures.

Abstract

World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.

Sources

Related papers