Learning Visual Feature-Based World Models via Residual Latent Action

arXiv:2605.07079 · cs.CV, cs.AI, cs.LG, cs.RO · Submitted 2026-05-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Learning Visual Feature-Based World Models via Residual Latent Action".

Jane: World models predict future transitions from observations and actions, and this work introduces Residual Latent Action (RLA) to create an efficient visual feature-based world model that outperforms existing methods.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to summarize the core thesis of "Learning Visual Feature-Based World Models via Residual Latent Action," they are addressing the challenges in visual feature-based world models that often result in blurry or collapsed predictions when using direct regression methods. The central claim is that a new latent action representation called Residual Latent Action, or RLA, can be learned simply from DINO residuals between frames.

Jane: That RLA is essentially a compact vector that captures the difference between two DINO tokens of consecutive frames and they found it has three key empirical properties: it’s predictive, generalizable to novel scenes and motion patterns even with limited data, and it encodes temporal progression.

Lu: Building on that RLA idea, the authors propose the RLA World Model, or RLA-WM. Instead of predicting the future DINO tokens directly, they predict these latent action values using flow matching based on current state and actions at a future time step as input conditions. Then they decode those predicted latent action values back into the actual visual features for that future frame.

Meng: So, instead of trying to solve a high-dimensional regression problem for the full image, they are solving a task of predicting this compact latent vector via flow matching, which sounds like a much more manageable mathematical problem.

Lalam: I see how that shift in focus makes the prediction process more structured and less prone to those kinds of artifacts, which is crucial when we're trying to build systems that need to interact reliably with the physical world.

Tom: And this approach is powerful because it allows them to achieve accurate future predictions while maintaining better computational efficiency than some other state-of-the-art feature-based and video diffusion world models.

Jane: They showed that RLA can be used for two distinct applications: first, learning a minimalist world action model from actionless demonstration videos, and second, training visual reinforcement learning entirely inside the RLA World Model using only offline videos to learn rewards.

Lu: The implication here is that this framework doesn't just sit there as a predictor; it’s a foundation for creating novel learning techniques, like policy learning from actionless video or visual reinforcement learning with video-aligned rewards.

Meng: That application aspect is where I get practical—if we can train RL purely inside this RLA-WM using offline data, that bypasses the need for constant online interactions or hand-crafted rewards, which simplifies the training pipeline significantly.

Lalam: And if it can improve performance on tasks like ManiSkill with robots like the XArm and UR10e, that shows a tangible path toward more capable robotic agents.

Conclusion: Tom: So, wrapping up this discussion on "Learning Visual Feature-Based World Models via Residual Latent Action," it’s clear the authors have successfully introduced a novel latent action representation that makes world modeling much more efficient while maintaining high accuracy in predicting future visual features.

Jane: The main implication is that by focusing on learning a compact residual latent action, they've provided a principled way to model the low-dimensional dynamics of visual transitions, moving us closer to models that can handle complex three dee manipulation <ref:2605.07079#pg0>.

Lu: What this means for the broader AI community is that we now have a tool—RLA and RLA-WM—that allows researchers to explore how predictive latent representations can be leveraged across different learning paradigms, from pure prediction to policy learning.

Meng: From an engineering standpoint, the practical implication is that we can deploy these models on devices where computational resources are limited because of their efficiency compared to older feature-based approaches.

Lalam: I think the real cultural impact comes from seeing how this structure enables more sophisticated AI agents that can learn complex skills just by observing demonstrations or interacting with simulated environments efficiently.

Tom: It really boils down to taking a complex visual prediction problem and simplifying it into predicting a compact latent action, which unlocks new ways for AI to learn and act in the world.

Rutgers University · Purdue University · University of Wisconsin-Madison

cs.CV, cs.AI, cs.LG, cs.RO

Submitted: 2026-05-08

Updated: 2026-10-05

Project page: https://mlzxy.github.io/rla-wm

Importance score: 91/100

The gist: World models predict future transitions from observations and actions, and this work introduces Residual Latent Action (RLA) to create an efficient visual feature-based world model that outperforms

Key concepts

Residual Latent Action (RLA)
RLA is a compact latent vector learned from the difference between DINO tokens of two frames. It is trained using a simple regression loss to reconstruct the future frame's tokens from the current ones. This representation is surprisingly predictive and generalizes well to new scenes.
RLA World Model (RLA-WM)
The RLA-WM predicts future visual features by first predicting the RLA values using flow matching, conditioned on current observations and actions. The final future state is then decoded from this predicted latent action and the current state, making the prediction process computationally efficient.
Flow Matching
Flow matching is a technique used in the RLA-WM to predict future states. It involves iteratively solving an ordinary differential equation to transform random noise into the desired latent action representation (RLA z), which guides the prediction of future visual features.
World Model-based RL (WMRL)
This framework trains reinforcement learning entirely within an RLA-WM learned from offline videos. It uses video-aligned rewards without requiring online interactions or handcrafted rewards, showing significant improvements in robot tasks like ManiSkill.

Terminology

Summary

World models predict future transitions from observations and actions, and this work introduces Residual Latent Action (RLA) to create an efficient visual feature-based world model that outperforms existing methods.

The gist

Residual Latent Action (RLA) is a compact latent action representation learned from DINO residuals, which, when used in the RLA World Model (RLA-WM), enables state-of-the-art prediction of future visual features by predicting RLA values via flow matching.

Residual Latent Action (RLA)

The core innovation is the Residual Latent Action (RLA), which encodes the residual between DINO tokens of two frames, denoted as st+h − st, into a compact latent vector, referred to as z. This representation is learned with a single regression loss to reconstruct st+h from st. Despite its simplicity, RLA exhibits three surprising empirical properties: (1) it is sufficiently predictive, allowing the decoder fdec to accurately reconstruct future DINO tokens in a single feedforward pass; (2) it generalizes to novel scenes and motion patterns, even when trained on limited data; and (3) the latent space exhibits a temporal topology, as interpolations between Gaussian noise and RLA yield results that approximate intermediate frames.

RLA World Model (RLA-WM)

Building upon RLA, the paper proposes the RLA World Model (RLA-WM), which predicts RLA values via flow matching. Instead of directly regressing DINO tokens st+h, the model first predicts RLA z via flow matching using st and actions at:t+h as input conditions. The final prediction of st+h is then achieved by decoding it from the predicted RLA z and the current state st. This approach significantly outperforms state-of-the-art feature-based and video diffusion world models on both simulation and real-world datasets, while remaining more efficient as the flow matching runs in the compact RLA space.

Applications of RLA

The framework enables two novel robot learning techniques:

  1. A minimalist world action model (WAM) that learns from actionless demonstration videos. This is achieved by extending a behavior cloning policy with a single linear layer that predicts RLA from the current observation, imposing no such coupling with heavy video generation backbones and adding no inference cost.

  2. A visual reinforcement learning (RL) framework trained entirely inside an RLA-WM learned from offline videos only, using a video-aligned reward and no online interactions or handcrafted rewards. This World Model-based RL (WMRL) yields a significant improvement on ManiSkill tasks for the XArm and UR10e robots.

Evaluation and Results

The RLA-WM was evaluated on the ManiSkill simulation suite and the IWS real-world dataset using metrics such as LPIPS, SSIM, DINO L1 distance, and FLOPs per inference. The results show that RLA-WM significantly outperforms all feature-based (DINO-WM, RAE, FM-WM) and video diffusion (Vid2World) baselines across all measured metrics on both ManiSkill and IWS. Furthermore, the model achieves high fidelity predictions with minimal hallucination and a computational efficiency second only to the direct regression of DINO-WM, demonstrating that RLA's compactness enables flow matching within a compact latent space. The WMRL framework also demonstrates gains over behavior cloning baselines, achieving success rates up to 63.6% on the Poke Cube task for the XArm robot.

Limitations and Future Directions

The authors identify four key limitations: (1) Task-irrelevant background motion can cause visual changes between st and st+h, suggesting a future direction to move from 2D image learning to 3D, projecting DINO tokens into 3D. (2) Visual changes may depend on s<t due to occlusion or partial observation, suggesting extending RLA to condition on multiple frames. (3) The model currently predicts only visual state evolution via RLA, not future proprioceptive states. (4) Evaluation focuses on small-scale datasets and simulation; scaling to internet-scale data in an open-world setting remains an open question. Additionally, the authors note that for certain robots like Panda, performance drops due to kinematic structure and limited action diversity in demonstration data.

How it works

The RLA autoencoder uses almost only self-attention and a single regression loss on st+h. The RLA World Model predicts future states by generating the residual latent action z, which is first predicted via flow matching using st and actions at:t+h as input conditions. This process involves iteratively solving an ODE to transform Gaussian noise into the final RLA z1. Finally, fdec decodes st+h from this predicted RLA z1 and st in a single feedforward pass.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, Learning Visual Feature-Based World Models via Residual Latent Action, and identified several high-impact areas for improvement in current AI systems.

Here are the specific improvements and what the resulting improved AI system can achieve:


)1. Improvement: Replacing Heavy Generative Pipelines with Compact Latent Dynamics Modeling (RLA-WM)

The core improvement is shifting world model prediction from high-dimensional pixel space (like video diffusion, which requires massive computation and suffers from hallucination) to a compact latent action space using the Residual Latent Action (RLA).

  • Specific Improvement: Implement a world model that predicts the residual latent action vector, rather than regressing raw future pixels or DINO tokens directly. This is achieved by learning RLA—the residual between sequential DINO tokens—and using this compact vector as a condition for flow matching to predict the next state.

  • What the Improved AI System Can Do:

The system can perform world modeling (future state prediction) significantly faster (orders of magnitude faster, as shown in Table 1) and with higher fidelity than current video diffusion models, especially in complex 3D interactions. It will be more reliable because it avoids the blurry or collapsed predictions inherent in direct regression. This allows for real-time or near-real-time simulation/planning.

)2. Improvement: Enabling Data-Efficient Policy Learning from Actionless Demonstrations (Minimalist World Action Model - WAM)

The RLA framework enables a novel way to train policies without requiring rich, labeled interaction data for every step.

  • Specific Improvement: Develop a Minimalist World Action Model (WAM) that extends standard Behavior Cloning (BC). This involves adding only a single linear layer to the existing policy backbone to predict the compact RLA vector. This eliminates the need for coupling with heavy video generation backbones during action prediction.

  • What the Improved AI System Can Do:

The system can learn robust control policies from actionless demonstration videos (where proprioceptive data is absent) with significantly higher success rates than current latent action learners (Table 2). This makes robot imitation learning scalable to large, unlabeled datasets where human demonstrations are scarce.

)3. Improvement: Visual Reinforcement Learning Entirely Inside the Learned Model (WMRL)

The RLA-WM framework allows for a fully autonomous RL loop that operates solely within the learned dynamics, bypassing the need for external simulators or handcrafted rewards during training.

  • Specific Improvement: Implement World Model-based Reinforcement Learning (WMRL). The system uses an offline demonstration to initialize its state and then performs rollouts entirely within the RLA-WM. Reward is defined via a Video Aligned Reward (VAR) based on DINO L1 distance between predicted and ground truth tokens.

  • What the Improved AI System Can Do:

The system can learn optimal control policies without any online interactions with the real world or manual reward engineering. This drastically improves sample efficiency for RL tasks (e.g., ManiSkill tasks like Push-T) and allows for massive, reproducible policy optimization across thousands of seeds, leading to superior performance over imitation learning (Table 3).

)4. Improvement: Achieving Cross-Embodiment Generalization in Latent Dynamics

The RLA representation is shown to be highly versatile, generalizing across different robot types and task complexities.

  • Specific Improvement: Utilize the learned RLA autoencoder, trained on a single dataset (ManiSkill), to predict dynamics for entirely unseen robot configurations (e.g., replacing a Panda arm with an XArm).

  • What the Improved AI System Can Do:

The system can rapidly adapt its world model to new robotic hardware or novel manipulation tasks with minimal fine-tuning, leveraging the generalized knowledge encoded in the RLA space. This reduces the need for task-specific world model training for every new robot setup.

Sources

Related papers