PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies

arXiv:2610.12285 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: I'm Rosa, and with me are Dev and Taro, guest researcher.

Dev: Today's paper: "PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies".

Rosa: The gist: PLaW-VLA introduces a framework that models task-relevant future states in a prediction-oriented representation space and conditions action generation on these predicted futures, enabling improved long-horizon control and generalization.

Dev: First, who's behind it and why it matters.

Paper summary: Rosa: Building on what we just heard, this PLaW-VLA paper introduces a framework that uses a mixture-of-transformers architecture to model task-relevant future states in a prediction-oriented representation space and conditions action generation on those predicted futures. The thesis is that learning to predict how the world evolves can provide vision-language action policies with predictive context for long horizon control, but its effectiveness depends on what future representation you model and how it conditions action generation.

Dev: Specifically, they claim that by predicting task-relevant future latent states instead of low-level visual reconstruction, they reduce the need to predict control-irrelevant visual details. They achieve this by using a frozen V-JEPA two encoder to define the space where those future features are predicted, rather than aiming for a reconstruction target space <ref:2610.12285#pg1>.

Taro: So the core claim is that this predictive context helps in long horizon manipulation under partial observability and execution perturbations because it models temporal context from observation history alongside anticipating future state evolution for action generation.

Rosa: Right, and they implement this by training two complementary objectives: one objective trains the latent world model to align its predictions with targets extracted from future observations using that frozen visual encoder, and another objective optimizes the action expert using a flowmatching objective over continuous action chunks.

Dev: And during robot policy learning, they jointly optimize both branches by stochastically dropping the predictive branch for each batch, which means they check if including those future latents actually improves the overall performance of the system.

Taro: I’m thinking about what this implies for autonomy—if an AI can learn to predict its future state in a structured way, it might be better at planning complex sequences than just reacting to the immediate visual input.

Rosa: It does seem like that, and on the RoboTwin Hard Horizon III benchmark, they report a +eleven point eight percentage-point gain over reactive policies <ref:2610.12285#pg2>. They also achieved seventy-two point seven percent task-weighted success on zero-shot LIBERO-Plus which is a plus for generalization under distribution shift in those tasks.

Dev: And from an engineering standpoint, the inference efficiency is quite good because they manage to avoid modeling low-level appearance details like texture and background, leading to about one/nineteen of the inference latency compared to generative world–action modeling at comparable policy performance <ref:2610.12285#pg2>.

Taro: So it’s not just about getting a higher score; it’s about finding a way for the system to handle those long sequences where you can't see everything upfront, which is where most current control methods struggle.

Rosa: That’s the point of PLaW-VLA: making sure the predictive context is actually task-relevant and that it conditions actions effectively. Now, let’s talk about what this means for the broader picture.

Conclusion: Dev: So to wrap up on PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies, the main point is that learning to predict world evolution can give VLA policies predictive context for long horizon control, provided you choose the right future representation and conditioning mechanism.

Rosa: The authors are Yu Liu, Hetian Guo, Tianlv Huang, Ziyi Cai, Wudi Chen, Hantang Wang1 through Wei Han1 and Peijun Tang2 who developed this work at Jilin University and other institutions.

Taro: It’s a solid piece of research because it shows that predictive supervision is important for learning informative future representations for the robot.

Dev: And the implication is that this approach to future state modeling allows the action expert to exploit those predictions effectively for better control, which leads to improved long-horizon control and generalization.

Rosa: It shifts the focus from just predicting pixels or images to modeling how things change, which is a significant way we teach robots what it means to have temporal context in complex tasks.

Taro: That's what matters for autonomy—it moves the capability beyond immediate perception into true state anticipation.

Dev: It’s a method that balances performance with efficiency, showing that you can get good control gains while keeping the model lightweight by focusing on dynamics over appearance.

Rosa: So PLaW-VLA is about using predictive latent world modeling to give vision-language action policies the context they need to handle complex, long sequences in manipulation tasks.

Yu Liu, Hetian Guo, Tianlv Huang, Ziyi Cai, Wudi Chen, Hantang Wang, Qiutong Liu, Yingzhi Peng, Wei Han

Jilin University 2Astribot Harbin Institute of Technology Shenzhen The Hong Kong Polytechnic University The University of Tokyo

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://rainyrobo.github.io/PLaW-VLA

The gist: The gist: PLaW-VLA introduces a framework that models task-relevant future states in a prediction-oriented representation space and conditions action generation on these predicted futures, enabling

Key concepts

Latent World Model (LWM)
This component predicts future latent states based on historical observations and task context, focusing on state transitions and interaction dynamics instead of low-level visual appearance. It operates in a prediction-oriented space to capture what the world will do next, which is crucial for planning long sequences of actions.
Prediction-Oriented Representation Space
Instead of modeling raw visual pixels, PLaW-VLA predicts future states within a representation space designed to emphasize task dynamics and state changes. This focuses the model on the relevant information needed for control, ignoring irrelevant visual noise like texture or background.
Structured Causal Attention
This mechanism allows predicted future latent states to directly condition the action generation model. It uses a causal attention structure to ensure that actions taken are informed by what is expected in the future, enabling better long-horizon control.

Terminology

Summary

The gist: PLaW-VLA introduces a framework that models task-relevant future states in a prediction-oriented representation space and conditions action generation on these predicted futures, enabling improved long-horizon control and generalization.

PLaW-VLA Framework

PLaW-VLA is a Vision-Language-Action (VLA) framework that integrates predictive latent world modeling with action generation Built on a Mixture-of-Transformers architecture, PLaW-VLA combines observation history, current task semantics, and predicted future latent states to generate continuous robot actions. The framework is designed around two key choices in future modeling for control. First, PLaW-VLA predicts task-relevant future latent states in a pretrained prediction-oriented representation space, emphasizing state transitions and interaction dynamics rather than low-level visual reconstruction. Second, these predicted future states directly condition the action model through structured causal attention.

Model Architecture

The PLaW-VLA architecture adopts a unified Mixture-of-Transformers (MoT) architecture consisting of three causally ordered experts corresponding to the factorization in Eq. (1). These experts include a vision-language expert for semantic grounding, a latent world model expert for predictive latent modeling, and an action expert for continuous action generation. The Vision-Language Model instantiates PaliGemma [32], which provides strong multimodal understanding and language-conditioned visual grounding capabilities. The Latent World Model predicts future latent trajectories conditioned on historical observations and task context, without explicitly modeling low-level visual appearance. Specifically, the latent world model predicts future latent states through the predictor Fz: zˆt+1:t+K = Fz(s1:K, zt−Ho:t, ϕt).

Training Objectives and Dynamics

PLaW-VLA utilizes two complementary objectives for training. The Latent Predictive Objective trains the LWM by aligning predicted future states with latent targets extracted from future observations using a frozen pretrained visual encoder, instantiated as V-JEPA 2 [31]. The Action Generation Objective optimizes the action expert with a flowmatching objective over continuous action chunks. During robot policy learning, the framework jointly optimizes future-state prediction and action generation by stochastically dropping the predictive branch for each batch This joint optimization is represented by Ljoint = LFM(D; Cm) + mλwmLwm(D), where m = 1 when using historical and predicted future latents, and m = 0 when omitting them.

Experimental Results

Experiments show strong performance across various benchmarks. On the RoboTwin Hard Horizon III benchmark, PLaW-VLA achieves a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III. Furthermore, it achieves 72.7% task-weighted success on zero-shot LIBERO-Plus. In real-world evaluation, PLaW-VLA outperforms both baselines on all four seen tasks, with the largest gains observed on Fold Towel and Pack Toy. Ablation studies confirm that predictive world modeling provides gains even without additional video pretraining.

Inference Efficiency

PLaW-VLA demonstrates a favorable performance–efficiency trade-off. Compared with Motus [16], PLaWVLA requires only about 1/19 of the inference latency at comparable LIBERO success. This efficiency is achieved by avoiding explicit modeling of low-level appearance details such as texture and background, using a lightweight latent world model in which multiple future states are predicted in parallel through future queries and directly used to condition the action expert without additional video generation or decoding.

The paper demonstrates that predictive context is crucial for long-horizon control by modeling task-relevant state evolution rather than low-level visual reconstruction. PLaW-VLA improves the performance–efficiency trade-off, achieving a +11.8 pp gain over π0.5 on RoboTwin Hard Horizon III, 72.7% average success on zero-shot LIBERO-Plus, and about 1/19 the inference latency of generative world–action modeling at comparable policy performance. The framework's success is supported by evidence showing that predictive supervision is important for learning informative future representations, while future conditioning allows the action expert to exploit them for control. Together, these findings highlight the importance of both what a VLA policy predicts and how those predictions inform its actions. The limitations noted include the coupling between prediction and closed-loop control through execution errors, and the underexplored scalability of the video-to-robot training paradigm.

--- Page 1 ---

PLaW-VLA: Predictive Latent World Modeling for Vision-Language-Action Policies Yu Liu1,2,∗ Hetian Guo1,∗ Tianlv Huang1 Ziyi Cai3 Wudi Chen1 Hantang Wang4 Qiutong Liu4 Yingzhi Peng5 Wei Han1 Peijun Tang2,† Jianan Wang2,† Zipei Fan1,‡ Zhiyuan Zha1 Xuan Song1 1 Jilin University 2Astribot 3Harbin Institute of Technology, Shenzhen 4The Hong Kong Polytechnic University 5The University of Tokyo https://rainyrobo.github.io/PLaW-VLA PLaW-VLA Policy Future Latent State Robot Action History Observation VLM WM Module Action Module Language instruction Current Observation Ego-Centric Real-Robotic Simulated-Robotic Long-horizon Fine-grained Deformable Simulation Figure 1: PLaW-VLA Videos and robot trajectories train a latent world model whose predicted future representations condition robot actions. We evaluate this approach on simulated and real-world manipulation, including multi-step and deformable-object tasks. Abstract: Learning to predict how the world evolves can provide vision-languageaction (VLA) policies with predictive context for long-horizon control, but its effectiveness depends on what future representation is modeled and how it conditions action generation. We introduce PLaW-VLA, which models task-relevant future states in a pretrained prediction-oriented representation space, reducing the need to predict control-irrelevant visual details. Built on a Mixture-of-Transformers architecture, PLaW-VLA conditions action generation on observation history, current task semantics, and predicted future states through structured causal attention. Experiments show a +11.8 percentage-point (pp) gain over reactive policies on RoboTwin Hard Horizon III and a +1.77 pp gain over reconstruction-oriented latent prediction on zero-shot LIBERO-Plus, supporting improved long-horizon control and generalization under distribution shift, respectively. By avoiding lowlevel visual reconstruction, PLaW-VLA lowers the burden of future prediction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world–action modeling at comparable policy performance. Keywords: Vision-Language-Action Models, Latent World Models

--- Page 2 ---

1 Introduction Vision-language-action (VLA) models adapt pretrained vision-language representations to languageconditioned robot manipulation [1, 2, 3, 4, 5, 6] Long-horizon manipulation requires not only interpreting the current scene but also tracking how the task evolves across successive interactions. Under partial observability and execution perturbations, the current observation alone may be insufficient to infer the underlying task state. Effective long-horizon control therefore benefits from both modeling temporal context from observation history and anticipating future state evolution for action generation. Recent studies incorporate predictive representations and future modeling into VLA through futurerepresentation alignment [7, 8], visual feature fusion [9], future-conditioned action generation [10, 11, 12], and joint world–action modeling [13, 14, 15]. This raises two key questions for control: what future representation should be predicted, and how should it condition action generation? Existing predictive targets derived from visual reconstruction or generation require modeling both task dynamics and appearance variations, even when the latter are irrelevant to control. This increases the prediction burden, while errors in these irrelevant components may propagate to action generation when the predicted future directly conditions the policy. We propose PLaW-VLA (Fig. 1), a VLA framework that integrates predictive latent world modeling with action generation. Built on a Mixture-of-Transformers architecture, PLaW-VLA combines observation history, current task semantics, and predicted future latent states to generate continuous robot actions. The framework is designed around two key choices in future modeling for control. First, PLaW-VLA predicts task-relevant future latent states in a pretrained prediction-oriented representation space, emphasizing state transitions and interaction dynamics rather than low-level visual reconstruction. Second, these predicted future states directly condition the action model through structured causal attention.

Improvements for AI systems

  1. Bold Header: Predictive Latent World Modeling for Long-Horizon Control

This system can perform long-horizon tasks by modeling task-relevant future states in a pretrained prediction-oriented representation space and conditioning action generation on these predicted future latent states through structured causal attention.

  1. Bold Header: Enhanced Generalization Under Distribution Shift

The framework supports improved generalization under distribution shift by focusing on predictive context for long-horizon control and demonstrating performance gains on tasks like LIBERO-Plus, suggesting that the model learns to anticipate state evolution rather than relying on low-level visual reconstruction.

  1. Bold Header: Efficient Inference with Reduced Latency

The system achieves a significant efficiency gain by "avoiding lowlevel visual reconstruction, enabling a lightweight latent world model with parallel future prediction and about 1/19 the inference latency of generative world–action modeling at comparable policy performance."

Sources

Related papers