Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation".
Dev: Token-World introduces an action-conditioned world model simulator that models dynamics directly in a compact, policy-aligned VLM visual-token space, avoiding intermediate RGB generation.
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: Let’s start by looking at the title and who put this work together; it's "Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation."
Dev: I see that title, Rosa; it clearly signals that the core innovation here is using those VLM tokens as the foundation for modeling dynamics, not just using them as input features.
Taro: It makes sense because when you’re dealing with complex manipulation tasks, the way a policy perceives the world through those tokens is often more direct than relying on predicted pixels.
Rosa: The authors are quite a team of researchers from different institutions; it shows this is coming from a broad cross-section of expertise in visual models and robotics.
Dev: That diversity is interesting because you need people who understand both the deep learning side, like those VLM features, and the practical side of how that affects control loops.
Taro: I think having researchers from different backgrounds helps you spot potential blind spots when designing a simulator, especially concerning how it interacts with real-world constraints later on.
Rosa: In simple terms, this paper is proposing a way to build a world model that doesn't waste time making pictures of the future; instead, it predicts the state directly in the compact token space that the robot's AI policy already understands.
Dev: That means when we evaluate an action sequence, we don't have to wait for an RGB image prediction pipeline to finish before checking if the action was good or bad.
Taro: It suggests a fundamental shift in how we think about simulation—moving away from simulating pixels and towards simulating the underlying semantic understanding of the world that drives the AI.
Rosa: It’s about making the simulator inherently policy-friendly from the start, which seems like a big step forward for practical robotics applications.
Dev: I agree; if we can reduce that simulation latency, it means we can run more iterations in a given training time or even have faster real-time response during planning.
Taro: And from an autonomy research angle, if the simulator is tightly coupled to the policy's input format, it might be better equipped to handle uncertainty and unexpected events within the learned world representation.
The paper's summary: Dev: So what’s actually happening inside Token-World then? It seems they are taking those high-dimensional VLM tokens and compressing them down into a much smaller set of compact representations, which they call ct.
Rosa: Right, they use a Semantic VAE to do that compression, mapping the large token space down to something much more manageable.
Taro: The paper mentions this is done while preserving the spatial arrangement of those tokens, which is key because you need the spatial relationship for manipulation planning.
Rosa: And once you have that compact state ct, they build their dynamics model—the transition model—on top of it using a flow-matching DiT architecture.
Dev: Flow matching is interesting; it’s a specific way to train the transition model by constructing a target flow and minimizing the loss function Lflow = E h w(tau) v theta - v* two two <ref:2610.00575#pg0>.
Taro: That training objective suggests they are aiming for very accurate future state predictions conditioned on history and action, which is essential for reliable rollouts.
Rosa: They define the "Compact VLM World State" as a combination of these compact visual tokens and proprioception, writing it as xt = (ct, pt), where pt represents proprioception or self-motion data.
Dev: That state is then fed into a spatio-temporal Transformer with factorized attention and a fixed-size GRU state to keep context going beyond the immediate attention window.
Taro: The GRU part seems like a smart addition for maintaining long-term context in those sequential predictions, which is something I’ve seen cause issues in simpler models during long tasks.
Rosa: So, the core summary is that they bypass the bottleneck of RGB generation by modeling action-conditioned dynamics directly within this compact token space, resulting in a simulator with lower latency.
Dev: That lower latency is what really matters for loop rates; if we can keep the step time down to zero point three five nine seconds as they report, it opens up possibilities for faster policy evaluation.
Taro: It’s about achieving better fidelity without sacrificing the speed needed for real-time planning when things get messy in an autonomous scenario.
The paper's improvements: Rosa: One of the big takeaways from this work is how they handle the representation itself, specifically by compressing those high-dimensional visual tokens into a compact representation ct = C phi(zt) in R N times d, where d is much smaller than the original dimension.
Dev: That compression step is crucial because it allows the dynamics model to learn in a lower-dimensional space, which simplifies things significantly for training and inference.
Taro: I wonder if that reduction to d=sixteen which they found to be the best overall trade-off, is generalizable across different types of manipulation tasks <ref:2610.00575#pg1>.
Rosa: Beyond just the compression, they use a specific training strategy where the VLM encoder and semantic VAE are frozen during world-model training; only the dynamics model is trained using that flow-matching objective.
Dev: That freezing part makes sense from a computational standpoint; you don't want to waste cycles re-training those large vision models just to learn how to simulate dynamics.
Taro: It implies that the learned visual representation quality is treated as a fixed, high-quality input feature for the dynamics learning phase, which simplifies things greatly for model development.
Rosa: They also introduced a way to handle inference where they only map the compact predicted tokens back to policy-facing features when the downstream policy actually needs them, using a mapping mechanism t+k = G phi(+k).
Dev: That selective reconstruction is smart because it avoids the cost of full RGB decoding on every single simulation step, which keeps the latency low during rollout.
Taro: This conditional reconstruction is what makes it feasible for long rollouts; you’re not wasting time generating things that the policy doesn't immediately use.
Rosa: In terms of evaluation, they showed that this direct simulation leads to better open-loop prediction fidelity compared to RGB-based baselines like IRASim and Ctrl-World.
Dev: And when we look at closed-loop evaluation, the correlation between simulated success rates and reference policy performance jumps from zero point five eight three up to zero point seven nine four when comparing it to Ctrl-World.
Taro: That jump in correlation tells us that this model is much more reliable for assessing how a policy will actually perform in the real world because the simulation outcome aligns better with the policy's actual success metrics.
Conclusion: Rosa: To wrap up, Token-World demonstrates that modeling action-conditioned dynamics directly in a compact VLM visual-token space avoids the bottleneck of intermediate RGB generation, leading to lower latency and better fidelity for robot manipulation tasks.
Dev: It seems the most significant practical win is that the closed-loop policy evaluation shows much stronger correlation with actual policy performance, which is vital for trustworthy simulation.
Taro: I think what this really means for autonomy is that we can build simulators that are faster and more accurate in terms of predicting how a learned policy will behave over long sequences.
Rosa: It suggests that focusing the world model on the representation interface used by the AI policy, rather than just simulating pixels, is a much more promising direction for scaling up these systems.
Dev: If we can maintain that low latency while improving fidelity, it opens doors for deploying these simulations in more complex and dynamic environments where real-world interaction is too slow or expensive.
Taro: For me, the implication is that future work should focus on how this compact token space generalizes when the environment presents truly novel physics or unexpected dynamics outside of the training set.
Rosa: So, we’ve talked about how Token-World uses direct VLM token simulation to improve speed and fidelity compared to previous RGB methods.
Dev: It really shows that for control engineers, reducing that simulation latency is a tangible gain in terms of loop rate and responsiveness.
Taro: And from an autonomy standpoint, the ability to get a correlation of zero point seven nine four on closed-loop metrics suggests this model can be used much more confidently when planning complex, history-dependent maneuvers.
Rosa: We’ve covered the title, summary, improvements, and what these results mean for real robotic applications using Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation.
Dev: It’s a compelling piece of work that moves simulation closer to being a truly policy-aligned tool.
Taro: I think we’re excited to see how this foundation can be used to test things that are currently too risky or time-consuming to try in the physical world.
Southern University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · Hong Kong University of Science and Technology 4 · Institute of Automation, Chinese Academy of Sciences · MUKA Robotics
cs.RO
Submitted: 2026-09-30
Updated: 2026-10-05
Comments: Submitted to IEEE International Conference on Robotics and Automation (ICRA) 2027
Project page: https://chuyaofu.github.io/Token-World
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 84/100
The gist: Token-World introduces an action-conditioned world model simulator that models dynamics directly in a compact, policy-aligned VLM visual-token space, avoiding intermediate RGB generation.
Key concepts
- Direct VLM Visual-Token Simulation
- Instead of predicting future RGB images, Token-World models robot dynamics directly in the compact visual-token space that a Vision-Language Model uses. This avoids an intermediate step where RGB images are generated and then converted back into policy inputs, making the simulation faster and more faithful to how the robot's AI actually perceives its environment.
- Compact VLM World State
- The state used for modeling is not raw high-dimensional tokens but a compact representation called 'ct', which is derived from the original VLM tokens. This compact visual token state, combined with proprioception (proprioception), forms the full world state, allowing the dynamics model to operate efficiently in a reduced space.
- Flow Matching Dynamics Modeling
- The transition model predicts the next compact state by using flow matching, a technique trained on a target flow. This method allows the simulator to learn how states evolve based on past history and actions within the compact token space, ensuring accurate predictions for future robot movements.
- S-VAE (Semantic VAE)
- A separately trained Semantic Variational Autoencoder is used to map high-dimensional VLM tokens into a compact Gaussian distribution. This process helps define the reduced state space, allowing the dynamics model to learn transitions in a smaller, more manageable representation while still retaining relevant information.
Terminology
Summary
Token-World introduces an action-conditioned world model simulator that models dynamics directly in a compact, policy-aligned VLM visual-token space, avoiding intermediate RGB generation. This approach addresses the bottleneck in existing world models where future RGB observations are predicted and then re-encoded into policy inputs. By compressing high-dimensional VLM features into a compact state, Token-World improves open-loop feature fidelity and closed-loop policy evaluation compared to current baselines while requiring lower simulation latency.
Core Innovation: Direct VLM Visual-Token Simulation
The central idea of Token-World is to simulate future dynamics directly in the compact VLM visual-token space
consumed by the downstream policy, rather than predicting future RGB images and re-encoding them. This design avoids an indirect interface between simulation and downstream policy execution,
which is a critical bottleneck in current world models. The paper investigates whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens.
This approach bypasses the need for intermediate RGB generation throughout the rollout,
leading to a simulator that operates on the representation interface used by VLA policies.
Token-World Architecture and State Representation
Token-World separates the policy representation from the world model's state space. Given policy-facing VLM tokens, it first maps each token into a compact representation: each token into a compact representation ct = Cϕ(zt) ∈ R N×d, d ≪ D0,
while preserving the spatial arrangement of tokens. The dynamics model is then learned in this reduced space. The Compact VLM World State
is defined as the combination of these compact visual tokens and proprioception: xt = (ct, pt).
This state is processed by a spatio-temporal Transformer with factorized spatial and causal temporal attention, supplemented by a fixed-size GRU state to maintain context beyond the finite attention window.
Dynamics Modeling via Flow Matching
The transition model predicts the next compact state conditioned on history and action: pθ (xt+1 x≤t, at).
The dynamics model is implemented as a flow-matching DiT, which is trained using an x0-parameterized flow objective with weighted shortcut forcing.
This training involves constructing a target flow where the transition model predicts the clean endpoint conditioned on available history. The resulting prediction induces a flow that is trained against the target flow using a loss function, specifically Lflow = E h w(τ)∥vθ − v∗∥2 2 i,
ensuring accurate future-state prediction in the compact space.
Training and Autoregressive Rollout
The training process involves several stages:
-
A separately trained Semantic VAE (S-VAE) is used to construct the compact state space, mapping high-dimensional VLM tokens into a diagonal Gaussian posterior qϕ(ct zt). The decoder Gϕ reconstructs the original policy-facing representation as z˜t = Gϕ(ct).
-
The dynamics model is trained using the flow-matching objective. The
VLM encoder and semantic VAE remain frozen throughout world-model training.
-
During inference (autoregressive rollout), future states are generated sequentially by
denoising an initialized noisy state conditioned on the current history and action.
Predicted compact visual states are mapped back to policy-facing features only when consumed by the downstream policy:zˆt+k = Gϕ(ˆct+k).
Evaluation and Key Results
Token-World is validated through three primary evaluation methods:
-
Open-loop prediction fidelity, which shows gains in VLM feature fidelity (e.g., cosine similarity) and policy-action consistency compared to RGB-based baselines like IRASim and Ctrl-World.
-
Closed-loop policy evaluation, where simulated success rates correlate more strongly with reference policy performance:
Token-World increases the Pearson correlation from 0.583 to 0.794 over Ctrl-World.
-
Simulation efficiency, achieving a
lowest per-step simulation time at 0.359 s,
yielding significant speedups over other simulators. Representation ablations confirm that the compact state design is crucial, with the dimension d=16 providing thebest overall trade-off for dynamics prediction.
The gist: Token-World models action-conditioned dynamics directly in a compact VLM visual-token space to improve open-loop fidelity and closed-loop policy evaluation while reducing simulation latency.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the Token-World paper, followed by a description of what these improved systems can achieve:
The core improvement lies in shifting world model simulation from an indirect, computationally expensive RGB-based pipeline to a direct, policy-aligned VLM visual-token space.
Here are the specific improvements and capabilities:
-
A world model simulator that predicts future states directly in the compact, action-conditioned VLM token space rather than predicting full RGB images and re-encoding them.
-
The implementation of a flow-matching Diffusion Transformer (DiT) for modeling dynamics within this compact token space, eliminating the need to rely on pre-trained videogeneration models for dynamics learning.
-
A
Compact Token Construction
module that compresses high-dimensional VLM visual tokens into a much smaller semantic token state (e.g., from 2560 channels down to 16 channels per token) while explicitly preserving the original spatial token layout to maintain policy compatibility. -
An autoregressive rollout mechanism that generates future states sequentially in the compact space using a spatiotemporal Transformer and a recurrent GRU state, ensuring long-horizon consistency without intermediate RGB generation during rollout.
-
A
Mapping Back
mechanism that only reconstructs the full VLM visual token representation from the compact predicted tokens when required by the downstream policy, completely bypassing expensive RGB reconstruction during simulation steps. -
A closed-loop policy evaluation framework that demonstrates a significantly stronger correlation between simulated success rates and reference policy performance (e.g., correlation of 0.794 vs 0.583 for baselines), leading to more reliable simulation of policy outcomes, especially over long rollouts where other models degrade.
-
A representation design strategy that shows the optimal compact state dimension for dynamics modeling is not monotonically increasing; it identifies a specific bottleneck dimension (e.g., 16 channels) that provides the best trade-off between reconstruction fidelity, diffusion modelability, and downstream dynamics prediction accuracy.
The resulting improved AI systems can achieve the following:
-
Maneuver complex, long-horizon robotic tasks with significantly higher success rates in closed-loop policy evaluation (e.g., reaching a target object or performing a sequence of manipulation steps) because the simulator maintains better fidelity over extended rollouts compared to existing models like Ctrl-World.
-
Execute complex planning and control policies with lower simulation latency (achieving 0.359 s per step), allowing for faster policy training and more efficient real-time interaction in robotic systems.
-
Generate highly consistent actions during imagined rollouts, as the dynamics model operates directly on the representation used by the VLA policy, ensuring that predicted states are perfectly aligned with what the downstream policy expects.
-
Perform robust offline evaluation of policies using simulated data derived from real-world trajectories, providing a more trustworthy metric for assessing generalizable robot skills than current methods that rely on indirect RGB reconstruction.
-
Develop novel representations for world models by understanding the trade-off between latent space smoothness (reconstruction fidelity) and the capacity to model complex dynamics (modelability), leading to more optimized compact state designs tailored specifically for robotics.
Sources
- World Models
- Dream to Control: Learning Behaviors by Latent Imagination
- WorldGym: World Model as An Environment for Policy Evaluation
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- WoW: Towards a World omniscient World model Through Embodied Interaction
- WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
- World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
- WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
- Octo: An Open-Source Generalist Robot Policy
- OpenVLA: An Open-Source Vision-Language-Action Model
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- WorldEval: World Model as Real-World Robot Policies Evaluator
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning
- DINO-Foresight: Looking into the Future with DINO
- Back to the Features: DINO as a Foundation for Video World Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- Revisiting Feature Prediction for Learning Visual Representations from Video
- Mask World Model: Predicting What Matters for Robust Robot Policy Learning
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving