MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling

arXiv:2610.12194 · cs.RO · Submitted 2026-10-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling".

Dev: The gist The introduction introduces MiniWAM, which instead predicts compact future representations learned from privileged current–future transitions,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: So we're talking about MiniWAM today. This paper is about how to make world models more efficient by ditching the need for those huge native future predictions.

Dev: Right. It focuses on learning compact future representations instead, which should cut down on training time significantly.

Taro: I’m curious if this compact representation still captures enough information to handle unexpected situations when the robot runs outside the lab.

Rosa: That's a big question for me, Taro. The core idea here is using something called PRISM, which combines inverse dynamics supervision with feature reconstruction to build these targets.

Dev: So they aren't just looking at what the future looks like in a standard visual space; they are training the model on privileged current-future transitions instead.

Taro: That makes sense, but if we only use those privileged transitions, how robust is it when things get messy or the world changes suddenly?

Rosa: Well, PRISM is designed to emphasize information that's actually relevant to the actions and what happens next in terms of control-relevant stuff.

Dev: And they achieve this by making sure the representation keeps useful future-state content even though it’s much smaller than a native prediction.

Taro: So, instead of predicting a massive image for every future step, we’re predicting a smaller target that still carries the necessary dynamics.

Rosa: Exactly. This is presented in the paper titled MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling, which introduces this concept to replace costly native future prediction with these compact predictive representations.

Dev: The paper outlines how they construct these targets using PRISM, which involves two main learning objectives during the first stage: inverse dynamics supervision and feature reconstruction.

Taro: What does that mean practically for the robot? Is it just making the robot move better, or is it about understanding *how* to move?

Rosa: It’s about understanding how to move because the inverse dynamics part biases those compact tokens toward behaviorally meaningful transition information related to actions.

Dev: And then they add feature reconstruction, which encourages those tokens to keep some of that predictive visual information without needing the full detail of the native representation.

Taro: I see. So it’s a trade-off: you want control relevance from dynamics, but you also want some fidelity to what the future actually looks like in terms of pixels.

Rosa: That’s right, and they show that this combination results in a representation that is substantially stronger in terms of behavioral structure compared to just using reconstruction alone.

Dev: The performance gains are pretty substantial on benchmarks like LIBERO-PLUS, where MiniWAM remains competitive with much larger world action models.

Taro: If it’s competitive, does that mean we don’t need those massive models anymore, or is it just a more efficient way to get them running?

Rosa: It suggests we can train policies using these compact targets and still achieve strong manipulation performance, even if the underlying model architecture is much smaller.

Dev: The numbers show an up to eight times speedup in world-action training when compared against native future-feature prediction methods, which is a big deal for practical deployment.

Taro: That speedup sounds very appealing for real-time control systems where latency matters a lot. What about the limitations they mention?

Rosa: The paper does point out that performance still degrades as the bottleneck of the model becomes too restrictive, and they found that their success rate drops from seventy-four point one percent at the largest bottleneck to sixty-one point seven percent at the smallest one.

Dev: So, a bigger model can handle this compression better than a smaller one if we're being strict about those constraints during training.

Taro: It sounds like there’s a trade-off between representation size and how well it handles complex dynamics or visual fidelity in those tight spaces.

Rosa: Exactly. The paper concludes that effective world action modeling doesn't require predicting the native visual future, and these compact learned representations offer a stronger and substantially more efficient target for policy learning.

Dev: So, to summarize MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling, they replace dense visual prediction with predictive targets learned through PRISM—combining inverse dynamics and reconstruction—to get faster training without losing control information.

Taro: It seems like the real implication for us is that we can scale our world models differently, focusing on structured transition information rather than just raw pixel matching.

Rosa: Right. This approach shows that scaling the video model isn't the only route to strong world-action learning; a compact model with a well-structured predictive objective can still be very competitive.

The paper's summary: Rosa: So, we’re looking at MiniWAM now, and what it really does is swap out those huge native future predictions for these much smaller predictive targets that they learn from just current transitions.

Dev: Right. Instead of trying to predict the full visual future step-by-step, MiniWAM focuses on learning a compact representation that captures the essential information needed for world modeling and action planning.

Taro: I’m wondering what this means for a robot when things get unexpected in the real world? Does this small target still hold enough information to handle weird situations?

Rosa: That’s the million-dollar question, Taro. The authors built their targets using something they call PRISM, which mixes inverse dynamics supervision with feature reconstruction.

Dev: So it’s not just guessing what the next frame looks like; they’re training the model to predict a state that is both behaviorally relevant for actions and visually consistent with the future.

Taro: So, when you look at their numbers, they say this compact approach is actually better than native prediction methods on benchmarks like LIBERO-PLUS. How much of a difference are we talking about?

Rosa: They report up to an eight times speedup in training time compared to what you get from native future prediction methods. That’s significant for getting models trained faster without sacrificing performance on those complex tasks.

Dev: The setup involves a unified objective function that balances the world model's prediction against the robot's actual actions, which helps guide both parts of the learning process simultaneously.

Taro: But what are the trade-offs? If you shrink that representation down, what kind of information do you lose about the future visual state?

Rosa: The reconstruction part of their training is specifically designed to make sure that even though it’s compact, it still preserves useful predictive information about what the future will look like.

Dev: And they show that this structure leads to a representation with substantially stronger behavioral structure when you compare it against methods that only use feature reconstruction.

Taro: That’s interesting because sometimes those smaller representations lose the fine details needed for complex interactions, but MiniWAM seems to find a way around that by focusing on the transition information.

Rosa: It really suggests that we don't need these massive, computationally expensive native future predictions to build strong world-action models.

Dev: This implies that we can get better performance and much faster training just by learning the right kind of compact target representation from existing transitions.

Taro: So, if I were a researcher looking at autonomy systems, this means we can design world models that are efficient and robust without needing to rely on the full power of massive visual prediction backbones.

Rosa: Exactly. The core idea is that the predictive objective itself needs to be structured in a specific way—using these transition-based targets—to get strong results.

Dev: And they’ve shown this approach works well under distribution shifts, meaning it holds up when the robot sees things it hasn't seen exactly before on benchmarks like LIBERO-PLUS.

Taro: It sounds like a much more practical path forward for deploying these models, because you aren't stuck needing massive computational resources just to get started.

Rosa: So we’ve looked at how MiniWAM uses PRISM to create these structured targets, and the numbers show it’s faster and more efficient than native prediction. Next up, we’ll talk about what this means for real-time deployment in a physical setting.

The paper's improvements: Rosa: So, we’re looking at what the authors suggest to improve this MiniWAM approach for future work and implications now.

Dev: Basically, they’re pointing out that you can keep pushing the efficiency gains by focusing on how those compact representations align better across different visual backbones.

Taro: So it sounds like the next step is making sure these compact targets are not just good for one specific model architecture, but actually generalize when we switch between things like DINOv3 and VAEs.

Rosa: That’s right, Taro. They show that their PRISM approach helps strengthen cross-backbone alignment even after you control those variables like the current task and current state confusion.

Dev: It means the structure they build into these targets is robust enough to bridge gaps between different visual encoders, which is something we need for more flexible AI systems.

Taro: So, if we want a world model that can handle any type of camera or sensor input without completely retraining the whole thing from scratch, this structural alignment in the target representation seems important.

Rosa: It is. And they also discuss how adding inverse dynamics supervision helps boost the success rate even when you make those models smaller, showing that it provides a real benefit beyond just feature reconstruction alone.

Dev: They highlight that by incorporating both objectives together, you get this complementary supervision that really guides the model toward better control-relevant behavior in a much tighter space.

Taro: That’s interesting because we always worry that if you compress the model too much, you lose the necessary feedback loop to actually execute good actions in a dynamic environment.

Rosa: Exactly. The implication is that scaling up doesn't always mean using bigger models; sometimes it means using a smarter way to structure the representation itself so it retains critical control information.

Dev: So, what does this change for us on the engineering side? It suggests we can design these compact targets with specific constraints in mind from the beginning, rather than just letting a large model try to figure out its own structure.

Taro: It shifts the focus from building a bigger visual model to designing a better predictive objective that captures dynamics and action relevance directly.

Rosa: Precisely. The paper suggests this path—focusing on structured transition information—is how we can get strong world-action learning without needing those massive native video prediction backbones for every single step.

Dev: This is important because it points toward a more efficient pipeline where the model learns to predict the compact target, and then the action policy just reads that structure efficiently.

Taro: So, we move from predicting pixels to predicting structured transition information, which seems like a much more stable way to build autonomy.

Rosa: That’s right. We'll be talking about how this compact world model actually runs in practice and what the latency implications are for real-world deployment next.

Conclusion: Rosa: So we’re wrapping up on MiniWAM, which boils down to using compact predictive targets learned from current transitions instead of predicting huge native future frames for world modeling and action planning.

Dev: It really shows that you can achieve strong performance without needing those massive native video prediction backbones, which is a big deal for deployment speed.

Taro: I think the biggest takeaway is that we can design models to focus on the transition information—the things that actually matter for control—instead of just trying to match every pixel in the future.

Rosa: Right. And they’ve shown that this compact approach is not just faster, but it also helps with cross-backbone alignment, meaning it works better even if you switch between different visual encoders.

Dev: That means our controllers can be more flexible because the underlying world model isn't tied to one specific vision architecture.

Taro: For autonomy research, this suggests that the focus should be on structuring those predictive targets in a way that inherently captures behavioral structure rather than just raw image reconstruction.

Rosa: Exactly. It’s about how you define that predictive objective—using inverse dynamics and reconstruction together—to get better results even when you shrink the target space.

Dev: I agree, and I think for my job, it means we can push the loop rate harder because the target representation is smaller and more manageable within our latency constraints.

Taro: It’s about creating a system that’s robust to distribution shifts by learning from those privileged transitions rather than trying to memorize every possible future state.

Rosa: So, in summary, MiniWAM proves that compact predictive representations can be a strong and substantially more efficient target for policy learning in world action modeling.

Dev: It’s a solid piece of work showing how to get high efficiency without sacrificing the quality of the learned world model.

Taro: It opens up a new direction where we design the representation structure first, instead of just relying on massive data and models to figure it out later.

Rosa: That’s right. We’ll be talking next about how this approach translates into actual deployment scenarios and what that means for real-world robotics.

Jie Chen, Ruofei Bai, Yuxin Cai, Yifeng Zhang, Chengyang He, Jun Li, Wei-Yun Yau

Department of Mechanical Engineering, National University of Singapore · Nanyang Technological University, Singapore

cs.RO

Submitted: 2026-10-08

Updated: 2026-10-08

Project page: https://j1dan.github.io/MiniWAM

The gist: The gist The introduction introduces MiniWAM, which instead predicts compact future representations learned from privileged current–future transitions, demonstrating that effective world–action

Key concepts

MiniWAM
A model designed to predict robot actions by learning compact future representations instead of predicting the full native visual future. It uses a two-stage process involving PRISM to create efficient targets for policy learning, leading to significant efficiency gains.
PRISM (Predictive Representations via Inverse Spatiotemporal Modeling)
A method used in Stage 1 to learn compact targets. It combines inverse-dynamics supervision, which focuses on action-relevant transition information, with feature reconstruction, which ensures the targets still capture useful future visual state content.
Compact Predictive Representation (Lt)
The small target representation learned by MiniWAM. Instead of predicting the entire native future image, this compact vector is much smaller and more efficient for policy learning while still containing critical information about how the world will change based on current observations.

Terminology

Summary

The gist The introduction introduces MiniWAM, which instead predicts compact future representations learned from privileged current–future transitions, demonstrating that effective world–action modeling does not require predicting native visual futures and that compact predictive representations provide a strong and substantially more efficient target for policy learning

MiniWAM Architecture

MiniWAM is designed to replace costly native future prediction with compact predictive representations learned from current–future transitions To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information In Stage 1, PRISM learns to encode current–future visual transitions into compact targets through inverse-dynamics supervision and feature reconstruction The transition encoder Eϕ is instantiated as a lightweight 3D convolutional network that maps privileged current–future features to a compact predictive representation Lt = Eϕ(xt, X+t), whose dimensionality is substantially smaller than that of the native future representation X+t After Stage 1, the PRISM encoder is frozen and Lt becomes the world-modeling target of MiniWAM, which jointly predicts the resulting PRISM target and robot actions from the current observation The unified objective for training is LFM = λLEw(τL)uˆL − u⋆ L∥2 + λaEw(τa)uˆa − u⋆ a∥2

PRISM Learning Objectives

PRISM learns the compact representation by combining two complementary objectives during Stage 1 Inverse-dynamics supervision encourages these tokens to retain transition information informative of the underlying actions, while feature reconstruction encourages them to preserve future-state content, without requiring the full detail of the native visual representation The Stage-1 objective is LS1 = λaEw(τa)uˆa − u⋆ a∥2 + λrecLrec + βDKL qϕ(Lt xt, X+t)∥ N (0, I) Inverse-dynamics therefore biases the compact state toward behaviorally meaningful transition information, while reconstruction preserves predictive information about the future visual state

Performance and Efficiency Gains

MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8× speedup in world–action training Empirically, MiniWAM reduces WAM training cost by up to 8× while outperforming native futurefeature prediction across simulation benchmarks A 0.25B MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks PRISM consistently reduces predictive co-training cost across both benchmarks and visual feature spaces, with larger gains under heavier prediction workloads At horizon 32, estimated total training speedup reaches 8.46× on ROBOTWIN 2.0, while Stage-2 training alone reaches 10.48×

Representation Analysis

PRISM contributes behavioral structure beyond feature reconstruction alone In the cross-backbone comparison, PRISM substantially strengthens cross-backbone alignment once task and current-state confounds are controlled Table 6 shows that PRISM induces substantially stronger behavioral structure

Ablation Studies

Adding inverse-dynamics supervision to reconstruction improves success rate by roughly 4.6 percentage points across all capacities, indicating that it provides complementary control-relevant supervision beyond feature reconstruction Performance degrades as the bottleneck becomes restrictive, with PRISM success rate decreasing from 74.1% at the largest bottleneck to 61.7% at the smallest

Cross-Backbone Alignment

PRISM substantially strengthens cross-backbone alignment once task and current-state confounds are controlled While global CKA provides little separation on the main view, PRISM markedly improves within-task and deconfounded alignment over both raw transitions and the architecture-matched random encoder The wrist view shows the same trend, with neighborhood agreement increasing from 0.11 to 0.33 and action RMSE decreasing from 0.37 to 0.27

Conclusion

These results demonstrate that effective world–action modeling does not require predicting the native visual future, and that compact predictive representations provide a strong and substantially more efficient target for policy learning MiniWAM retains strong performance under distribution shifts in LIBERO-PLUS, achieving 73.6% success and outperforming Fast-WAM while approaching DreamWAM, despite using roughly an order of magnitude fewer world–action model parameters and no pretrained video-generation backbone A compact model can remain highly competitive when its predictive objective is defined over an appropriately structured future representation This demonstrates that scaling the video model is not the only route to strong world–action learning: a compact model can remain highly competitive when its predictive objective is defined over an appropriately structured future representation The project page is available at https://j1dan.github.io/MiniWAM

AI Use Statement

We used generative AI tools to assist with refinement of experimental and analysis protocols, code implementation and debugging, interpretation of experimental results, and polishing of manuscript text

Ethics Statement

This work studies efficient robot policy learning using existing datasets and simulation benchmarks Performance on these benchmarks does not establish safety or reliability in real-world environments, where distribution shifts and policy errors may cause physical harm Deployment on physical robots requires task-specific safety validation, appropriate operational constraints, and human oversight Use and redistribution of the datasets and pretrained models remain subject to their respective licenses and usage terms Our planned code release is intended to support transparent evaluation and further research into reliable robot learning

Reproducibility Statement

The main paper and appendix describe the method and experimental protocols Section 3 specifies the two-stage learning procedure, while Appendices A–D provide architecture details, training hyperparameters, dataset and evaluation settings Appendix E details the deterministic sampling, representation readouts, and analysis procedures We will publicly release the implementation, experiment configurations, trained model checkpoints, and evaluation scripts upon acceptance to facilitate reproduction of our results

References

Kevin Black et al. π0.5: a vision-language-action model with open-world generalization Kevin Black et al. π0: A vision-language-action flow model for general robot control Qingwen Bu et al. Learning to Act Anywhere with Task-centric Latent Actions Jisong Cai et al. Aha-wam: Asynchronous horizon-adaptive worldaction modeling with observation-guided context routing Rui Cai et al. Xiaomi-robotics-0: An open-sourced vision-language-action model with real-time execution Jun Cen et al. Worldvla: Towards autoregressive action world model Tianxing Chen et al. Robotwin 2.0: A scalable data generator and benchmark with strong domain randomization for robust bimanual robotic manipulation StarVLA Community. Starvla: A lego-like codebase for vision-language-action model developing Senyu Fei et al.

Improvements for AI systems

  1. Bold header: Compact World Modeling for Efficient Training

MiniWAM can achieve up to an 8× speedup in world–action training and be competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks by replacing costly native future prediction with compact predictive representations learned via PRISM.

  1. Bold header: Enhanced Control-Relevant Representation Learning

The PRISM encoder learns a representation that preserves information about both the actions underlying an observed transition and the resulting visual state, which leads to substantially stronger behavioral structure compared to reconstruction-only methods, as shown by the increase in action RMSE reduction from 0.414 to 0.356 when adding inverse dynamics supervision on DINOv3 bottlenecks.

  1. Bold header: Cross-Backbone Generalization

The PRISM representation substantially strengthens cross-backbone alignment, meaning it reorganizes native visual features at both global and local scales in a way that is more consistent across fundamentally different visual backbones like VAE and DINOv3.

  1. Bold header: Robustness to Distribution Shifts

MiniWAM, trained with PRISM targets, shows improved performance under distribution shifts in LIBERO-PLUS, achieving 73.6% success compared to native co-training's 50.2%, demonstrating that compact predictive representations provide a strong and substantially more efficient target for policy learning.

  1. Bold header: Efficient Inference and Deployment

The final inference system allows deployment with only the current observation, instruction, and proprioception, initializing latent states and jointly denoise them through the world and action experts to execute actions in a receding-horizon manner, requiring no need for future observations at deployment time.

Sources

Related papers