GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions".
Dev: The gist The paper introduces GeniWorld, an interactive world model that generalizes robustly across unseen scenarios by conditioning on visual actions to enable closed-loop interaction with policies and human operators.
Rosa: First, who's behind it and why it matters.
Paper summary: Rosa: We’ve been looking at GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions. To recap, the thesis of this paper is that they are introducing a generalizable interactive world model conditioned on embodied visual actions to enable closed-loop interaction with both robot policies and human operators.
Dev: The main claim is that even with limited demonstrations, this model generalizes robustly across diverse scenarios and allows for the generation of novel behaviors.
Taro: They achieve this by using an embodiment-specific kinematic model first to convert numerical action sequences into dense robot motion sequences, which then feeds into a pretrained video generative model to encode visual actions and scene observations into spatially aligned latent representations.
Rosa: The methodology involves building an autoregressive model with causal attention, ensuring future predictions depend strictly on current robot actions and historical states, which is crucial for closed-loop interaction.
Dev: They use URDF rendering to convert those robot actions into visual motions, then these are encoded and concatenated with noisy video latents to form a combined representation that feeds into a causal DiT that predicts future videos via flow matching.
Taro: The training objective uses flow matching to predict the next observation conditioned on the history and the corresponding action token, which is formulated as L = E t,s,z t+one epsilon v theta(z)(s) t+one s, z t a t+one c - z(s) t+one <ref:2608.06332#pg1>.
Rosa: Essentially, they are showing how this setup allows for precise spatial guidance for manipulation and captures fine-grained interaction details through visual action conditioning.
Dev: They also built in KV caching during inference to maintain high-quality video generation while enabling that closed-loop interaction with the robot policy or human operator.
Taro: So, it’s about creating a system where the model can actually interact dynamically, not just passively predict what will happen next.
Rosa: And they claim that even with limited demonstrations, this setup allows for generalization to out-of-distribution scenarios and enables the generation of novel behaviors.
Dev: The paper summarizes their key contributions as introducing GeniWorld, which learns from fixed scene-specific demonstrations and generalizes robustly to diverse unseen scenarios, supporting robot manipulation in rich synthesized imagination spaces.
Taro: Plus, they propose an autoregressive generative model conditioned on visual actions to decouple embodiment motion from scene dynamics, enabling explicit interaction modeling and closed-loop interaction with both robot policies and human teleoperators.
Rosa: And finally, they demonstrate that GeniWorld serves as a robust policy evaluator and enables the synthesis of rich manipulation data from limited real-world demonstrations that enhances downstream policy performance across diverse robotic systems.
Conclusion: Rosa: So, wrapping up on GeniWorld: A Generalizable Interactive World Model for Robotic Manipulation via Visual Actions. The title itself suggests a model that’s not just about predicting what happens, but about active interaction.
Dev: And the authors are Chenghao Gu and team, so they focused on making sure this model works reliably in real-world robotic manipulation tasks.
Taro: What this really means for the field is that it provides a powerful framework for scalable imagination spaces where you can generate data from limited demonstrations without needing massive manual scene construction.
Rosa: It moves the focus toward using visual actions as a way to precisely guide embodiment motion, which should help in developing more controllable and interactive embodied AI systems.
Dev: They show that this model can effectively synthesize rich and varied manipulation data that significantly boosts downstream policy performance under novel layouts and in complex visual environments.
Taro: It’s about moving toward models that are robust enough to handle the messy reality of real-world interaction without requiring perfect prior knowledge of every single possible environment.
Rosa: GeniWorld offers a way to move beyond just learning from demonstrations to creating synthetic data that helps policies learn better in complex, unseen situations.
Shenzhen International Graduate School, Tsinghua University
cs.RO
Submitted: 2026-08-06
Updated: 2026-10-08
Project page: https://chenghaogu.github.io/GeniWorld/Interactive
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: The gist The paper introduces GeniWorld, an interactive world model that generalizes robustly across unseen scenarios by conditioning on visual actions to enable closed-loop interaction with policies
Key concepts
- Autoregressive Robotic World Model
- This is a model that predicts future states or observations sequentially, where each prediction depends on previous ones. GeniWorld uses this structure to transform robot actions into visual action representations, allowing it to simulate complex interactions step-by-step.
- Visual Action Representations
- Instead of using simple numerical actions, GeniWorld converts robot movements into visual tokens. This allows the model to understand the spatial and temporal context of an action visually, which is crucial for grounding control in a physical scene.
- Flow Matching Training Objective
- This is a specific training method used to teach the model how to predict future observations. It guides the model by matching its predictions against real data using a continuous flow, ensuring high-fidelity generation of novel scenarios.
Terminology
Summary
The gist The paper introduces GeniWorld, an interactive world model that generalizes robustly across unseen scenarios by conditioning on visual actions to enable closed-loop interaction with policies and human operators.
How it works
-
GeniWorld is an autoregressive robotic world model that transforms action inputs into visual action representations, enabling closed-loop interaction with human operators and robot policies.
-
It builds on pretrained video generative models and uses URDF-based rendering to transform numerical actions into visual action representations, which enables spatially grounded action control.
-
The model encodes visual actions and scene observations into spatially aligned latent representations by concatenating them, forming a joint representation z = [zv; za] in R 2C×L×H′×W′.
-
It constructs an autoregressive model with causal attention, ensuring that future predictions depend strictly on current robot actions and historical states.
-
The training objective uses flow matching to predict the next observation conditioned on the history and the corresponding action token, formulated as L = Et,s,zt+1,ϵ vθ(z)(s)t+1, s, z≤t za,t+1 c − z˙(s)t+1 22.
Key Contributions
** We introduce GeniWorld, an interactive world model that learns from fixed-scene demonstrations and generalizes robustly to diverse unseen scenarios, supporting robot manipulation in rich synthesized imagination spaces. **
** We propose an autoregressive generative model conditioned on visual actions to decouple embodiment motion from scene dynamics, enabling explicit interaction modeling and closed-loop interaction with both robot policies and human teleoperators. **
** We demonstrate that GeniWorld serves as a robust policy evaluator and enables the synthesis of rich, varied manipulation data from limited real-world demonstrations, thereby enhancing downstream policy performance across diverse robotic systems. **
Experimental Evaluation
-
GeniWorld achieves superior generative performance over the baselines and demonstrates strong zero-shot interaction modeling across OOD scenarios.
-
Across multiple real-world manipulation tasks, policy success rates measured in GeniWorld correlate positively with real-world performance, supporting its utility as a scalable policy evaluator.
-
Given a limited set of real demonstrations, GeniWorld generates rich and varied manipulation data that improve downstream policy performance under novel layouts and in complex visual environments.
-
In the Clean-to-Random setting, GeniWorld maintains remarkable fidelity, retaining low FID and FVD scores of 13.08 and 20.15, demonstrating superior generation quality even when scene appearance, object instances, object placements, and task-relevant spatial layouts differ significantly from those seen during training.
-
Visual action conditioning consistently maintains superior generation quality with fewer denoising steps than the other action representations.
Policy Improvement
-
Policies are fine-tuned from the π0 vision-language-action (VLA) model using the collected real-world demonstrations, initialized with the same first-frame visual observation in GeniWorld and the realworld environment.
-
The policy is trained using the conditional flow-matching objective LFM = Eh∥vθ(Aτt, ot) − (At − ϵ)∥22.
-
Policies augmented with GeniWorld-synthesized data demonstrate superior performance across diverse OOD conditions, such as spatial rearrangement, novel instances, distractors, and lighting shifts.
-
Combining both Spatial-Gen and Diverse-Gen sources performs best across all settings, reaching 70.0% under spatial rearrangement, 72.5% with distractors, 70.0% on novel instances, and 52.5% under lighting shifts.
-
These results demonstrate that GeniWorld effectively mitigates data scarcity in robot policy learning by synthesizing highly diverse data without the high cost of manual scene construction and data collection.
The work highlights the strong potential of generalizable world models to provide scalable imagination spaces for embodied robot learning.
--- Page 1 ---
GeniWorld is an autoregressive robotic world model that transforms action inputs into visual action representations, enabling closed-loop interaction with human operators and robot policies.
Trained on limited scene-specific demonstrations, GeniWorld generalizes to out-of-distribution (OOD) scenarios to produce high-fidelity observation predictions.
Furthermore, it synthesizes rich manipulation trajectories to boost downstream policy performance and robustness under diverse conditions.
--- Page 2 ---
Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution (OOD) scenarios.
We argue that directly conditioning the world model on robot motion provides precise spatial guidance for manipulation and captures finegrained interaction details.
Meanwhile, conditioning scene dynamics on robot motion more directly reflects the physical changes driven by robot–environment interactions.
--- Page 3 ---
For the target robotic system, an embodiment-specific kinematic model first converts numerical action sequences into dense robot motion sequences.
Building on a pretrained video generative model, GeniWorld encodes visual actions and scene observations into spatially aligned latent representations.
By explicitly decoupling embodiment kinematics from environmental dynamics, our model mitigates scene overfitting and facilitates modeling of robot-environment interactions.
--- Page 4 ---
First, GeniWorld achieves superior generative performance over the baselines and demonstrates strong zero-shot interaction modeling across OOD scenarios.
Second, across multiple real-world manipulation tasks, policy success rates measured in GeniWorld correlate positively with real-world performance, supporting its utility as a scalable policy evaluator.
Third, given a limited set of real demonstrations, GeniWorld generates rich and varied manipulation data that improve downstream policy performance under novel layouts and in complex visual environments.
--- Page 5 ---
We convert robot actions into visual motions via URDF rendering. These motions are subsequently encoded into latent representations and concatenated with noisy video latents. The combined representation serves as input to a causal DiT that predicts future videos via flow matching.
During inference, the initial scene image is provided as the first frame. Actions are passed through the URDF-based renderer to produce motion conditions, and the model predicts future observations while leveraging KV caching to maintain historical context.
--- Page 6 ---
We conduct comprehensive experiments to evaluate GeniWorld from three perspectives.
First, GeniWorld achieves superior generative performance over the baselines and demonstrates strong zero-shot interaction modeling across OOD scenarios.
Second, across multiple real-world manipulation tasks, policy success rates measured in GeniWorld correlate positively with real-world performance, supporting its utility as a scalable policy evaluator.
Third, given a limited set of real demonstrations, GeniWorld generates rich and varied manipulation data that improve downstream policy performance under novel layouts and in complex visual environments.
--- Page 7 ---
We construct our world model based on a video diffusion backbone [54]. As shown in Fig. 2, visual actions are processed via a causal 3D VAE encoder into action latents za ∈ R C×L×H′×W′.
This design aligns the embodiment dynamics with the pretrained visual latent space, ensuring strict spatial correspondence between each position in the action and scene latents.
By leveraging this action-embedding mechanism, we introduce minimal architectural modifications to the pretrained video backbone, thereby preserving its rich generative priors for interaction modeling [9].
--- Page 8 ---
Table I shows that Ours w/ numerical actions achieves the best results in Clean-to-Clean and Clean-to-Random settings.
In the in-domain evaluation, our model consistently achieves superior performance across all six metrics.
Compared to ablation baselines using numerical action control or alternative explicit representations (e.g., end-effector pose and skeleton), our model demonstrates significantly higher alignment with ground-truth observations.
Furthermore, under domain shifts from clean to randomized environments, prior works such as IRASim experience severe degradation, with FID and FVD dropping to 174.52 and 191.26, respectively.
In contrast, GeniWorld maintains remarkable fidelity, retaining low FID and FVD scores of 13.08 and 20.15.
--- Page 9 ---
Fig. 3 shows that GeniWorld preserves the commanded robot motion and produces interaction outcomes that closely match the ground truth, whereas competing methods fail to generate the expected results in out-of-distribution (OOD) scenarios.
Furthermore, the ablation results show that our dense visual actions enable more accurate interaction modeling.
--- Page 10 ---
Fig. 5 shows that visual action conditioning consistently maintains superior generation quality with fewer denoising steps than the other action representations.
--- Page 11 ---
Fig.
Improvements for AI systems
-
textbfGeneralizable Interactive World Model for Robot Manipulation (GeniWorld): Development of a robust, closed-loop system that transforms numerical actions into visual action representations to enable
closed-loop interaction with human operators and robot policies.
This allows for scalable policy learning and evaluation across diverse real-world environments, even inout-of-distribution (OOD) scenarios.
-
textbfExplicit Decoupling of Embodiment Kinematics from Scene Dynamics: The model achieves this by mapping
numerical actions at+1:t+H through the robot’s URDF model and forward-kinematics system and rendering the corresponding embodiment motion,
whichmitigates scene overfitting
by ensuring that conditioning is onembodied visual actions
rather than low-dimensional action vectors entangled with environmental changes. -
textbfRobust Policy Evaluation Under Environmental Perturbations: GeniWorld serves as a reliable policy evaluator because it
remains robust in noisy conditions,
demonstratingstrong positive correlation between real-world and simulated performance
even when subjected toenvironmental perturbations
such as visual distractors, spatial rearrangement, or lighting shifts. -
textbf Scalable Data Synthesis for Policy Improvement: The system can generate diverse manipulation data from limited demonstrations by using advanced image-generation models to create
high-variance manipulation scenarios through instruction-driven editing,
allowing policies to be trained under regimes likeReal + Spatial-Gen + Diverse-Gen
to achieve gains up to 72.5% in performance across challenging OOD conditions. -
textbf Efficient Inference for Interactive Control: The model maintains high interactivity by leveraging KV caching and achieving a
∼ 10× inference speedup with an FVD degradation of only ∼ 2%
when using visual action conditioning, whicheffectively increases the interaction rate
compared to numerical-action-conditioned baselines.
Abstract
Generalist robot policies exhibit strong capabilities, but their robustness in complex and unseen environments remains limited. Scaling robot learning and evaluation in diverse real-world environments remains costly and challenging. Action-conditioned world models offer a promising alternative, but they often suffer from limited action controllability and poor generalization to out-of-distribution (OOD) scenarios. To this end, we present GeniWorld, an interactive world model for robots that generalizes robustly across unseen scenarios. Building on pretrained video generative models, we use URDF-based rendering to transform numerical actions into visual action representations, enabling spatially grounded action control. By explicitly decoupling embodiment kinematics from environmental dynamics, our model mitigates scene overfitting and facilitates modeling of robot-environment interactions. To achieve closed-loop control, we construct an autoregressive video prediction model integrated with high-frequency robot kinematic control, enabling interaction with both robot policies and human teleoperators. In our experiments, even when trained solely on limited fixed-scene data, our model achieves superior in-domain performance and robust zero-shot generalization to highly randomized, unseen environments. For downstream applications, GeniWorld serves as a scalable policy evaluator that remains reliable under environmental perturbations. Furthermore, even with limited real-world demonstrations, GeniWorld generates diverse manipulation trajectories within the world model, improving downstream policy performance and robustness in complex environments.
Sources
- $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control
- $\pi^{*}_{0.6}$: a VLA That Learns From Experience
- ${\pi}_{0.7}$: a Steerable Generalist Robotic Foundation Model with Emergent Capabilities
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
- Gemini Robotics: Bringing AI into the Physical World
- Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
- Causal World Modeling for Robot Control
- World Action Models are Zero-shot Policies
- Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
- Fast-WAM: Do World Action Models Need Test-time Future Imagination?
- LIBERO-PRO: Towards Robust and Fair Evaluation of Vision-Language-Action Models Beyond Memorization
- LIBERO-Plus: In-depth Robustness Analysis of Vision-Language-Action Models
- RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
- MaskWAM: Unifying Mask Prompting and Prediction for World-Action Models
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- WorldEval: World Model as Real-World Robot Policies Evaluator
- WorldGym: World Model as An Environment for Policy Evaluation
- Interactive World Simulator for Robot Policy Training and Evaluation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving