RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation

arXiv:2607.06558 · cs.RO · Submitted 2026-07-07 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation".

Dev: RynnWorld-Teleop is a generative digital teleoperation framework that decouples data collection from physical constraints by replacing a real robot with a generative world model,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: Well, the title itself, "RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation," suggests a system where the world model is actively conditioned by the actions of an operator. It sounds like they are focusing on making the synthetic world directly responsive to human input in a controlled way.

Dev: And looking at the authors, I see a mix of expertise here, which is usually good for complex systems work; you have people from different labs contributing to this digital teleoperation approach.

Taro: The core implication I see immediately is that if this works as described, it means we can train policies based on human demonstrations without needing thousands of hours of physical robot time to gather that data. That's a huge shift in how we think about imitation learning for complex tasks.

Rosa: Precisely, Taro; the paper suggests that by using an embodiment-agnostic action label derived from the hand pose stream, we get trajectories that can be retargeted to any target robot later on. It’s about making the data reusable across different hardware setups.

Dev: The implication for deployment is significant because it bypasses the need for every single demonstration to be tied to a specific physical robot and a fixed workspace, which is where so much of our current data collection bottleneck lies.

Taro: That decoupling addresses the scalability issue directly, suggesting that we can gather massive amounts of data purely from operator imagination rather than being limited by hardware availability.

The paper's summary: Rosa: So, to summarize what the paper describes, RynnWorld-Teleop is this framework where an operator's hand pose stream drives a robot-centric generative world model to synthesize high-fidelity egocentric videos from just one reference image. It’s essentially digital teleoperation that creates the necessary data on demand.

Dev: That means they are using a system built around a three dee Variational Autoencoder and a Transformer denoiser, conditioned by both the reference image latent and a depth-aware skeletal control latent. It sounds like they're building this world model piece by piece to predict the velocity field.

Taro: The methodology seems to involve rendering twenty-one-joint hand poses with camera-distance-modulated color and radius to give explicit three dee cues, which then gets projected into a latent space for control. That's a clever way to bridge that gap between what the operator sees and what the robot needs to do.

Rosa: And they use this aligned control latent in an additive patch-embedding scheme, using distribution alignment techniques to ensure the video generation stays close to what's expected from those action signals. It’s about making sure the synthesized output matches the intended movement precisely.

Dev: The training process itself is also structured in stages: first pretraining on large egocentric human videos for fundamental dynamics, and then fine-tuning on paired human–robot data to bridge that embodiment gap through inverse kinematics mapping.

The paper's improvements: Rosa: Regarding the improvements they propose, the key ones focus on making the action representation robust. They introduced depth-aware skeletal conditioning specifically to resolve that ambiguity between 2D projections and actual three dee dynamics, which is a big step.

Dev: I’m interested in their use of streaming autoregressive distillation; it's designed to convert a slow, bidirectional teacher model into a causal student model that can handle real-time interaction efficiently. That seems like they are trying to solve the latency problem inherent in complex generative models.

Taro: The progressive cross-domain training strategy is another important improvement because it allows the system to first absorb general manipulation priors from human videos before focusing on mapping those gestures specifically onto robotic actions via inverse kinematics.

Rosa: And they mentioned mitigating drift through a technique called Chunked Re-anchoring, where they re-anchor the generation process by providing the actual egocentric frame from the robot's camera at each subsequent chunk. This sounds like a practical fix for maintaining consistency in long generations.

Dev: That chunking and re-anchoring must be carefully managed because if that re-anchoring introduces any significant jitter or delay, it could completely ruin the control loop stability we need for real-time application.

Conclusion: Rosa: So, to wrap up on this paper, RynnWorld-Teleop presents a complete digital teleoperation system where raw operator motion is converted into paired video and action trajectories through retargeting and skeletal-conditioned synthesis. The main implication is that it serves as a high-fidelity data engine that both substitutes for and amplifies physical teleoperation in robot learning.

Dev: I see the results showing zero-shot transfer to real robots with success rates up to one hundred percent for tasks like Block Pushing and Bimanual Lifting when policies are trained purely on RynnWorld-Teleop data. That's a very strong indicator of its potential for practical deployment in complex manipulation scenarios.

Taro: I think the consistent overlap between the synthetic data distribution and real-world trajectories shown via t-SNE analysis really supports the idea that this model is capturing the underlying distribution of robotic manipulation effectively, which is what we need for generalization.

Rosa: It sounds like a very promising direction for scaling up robot learning by making it independent of physical hardware limitations. We have a lot to think about as we look at how this digital teleoperation paradigm can be applied across different domains and tasks in the future.

DAMO Academy, Alibaba Group

cs.RO

Submitted: 2026-07-07

Updated: 2026-09-29

Comments: Project Page: https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.io, Github: https://github.com/alibaba-damo-academy/RynnWorld-Teleop

Code: https://github.com/alibaba-damo-academy/RynnWorld-Teleophttps:

Project page: https://alibaba-damo-academy.github.io/RynnWorld-Teleop.github.iohttps://github.com/alibaba-damo-academy/RynnWorld-Teleophttps://huggingface.co/Alibaba-DAMO-Academy/RynnWorld-Teleophttps://www.modelscope.cn/models/DAMO_Academy/RynnWorld-Teleop

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 91/100

The gist: RynnWorld-Teleop is a generative digital teleoperation framework that decouples data collection from physical constraints by replacing a real robot with a generative world model, allowing operator

Key concepts

RynnWorld-Teleop
A generative digital teleoperation framework that uses an action-conditioned world model. It allows an operator's hand pose stream to drive a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image, creating data on demand.
Embodiment-agnostic action label
A method used to derive an action label from the hand pose stream. This label is designed to be reusable across different hardware setups, allowing trajectories generated by the model to be retargeted to any target robot later on.
Depth-aware skeletal conditioning
A technique introduced in the paper that adds depth-aware skeletal conditioning. This helps resolve ambiguity between 2D projections and actual three-dimensional dynamics, making the action representation more robust for control.

Terminology

Summary

RynnWorld-Teleop is a generative digital teleoperation framework that decouples data collection from physical constraints by replacing a real robot with a generative world model, allowing operator hand-pose streams to drive a robot-centric generative world model to synthesize high-fidelity egocentric videos from a single reference image. This paradigm is called digital teleoperation. The recorded pose stream serves as an embodiment-agnostic action label transferable to any target robot via standard retargeting, yielding complete state-action trajectories for imitation learning independent of physical hardware. RynnWorld-Teleop integrates depth-aware skeletal conditioning, progressive human-torobot training on a video Diffusion Transformer (DiT), and streaming autoregressive distillation into a single-pass inference pipeline, enabling 40+ FPS real-time interactive generation on a single H100 GPU. Policies trained exclusively on RynnWorld-Teleopgenerated data achieve effective zero-shot Sim2Real transfer across dexterous and diverse bimanual tasks, and augmenting real-world datasets with digitally teleoperated data consistently improves success rates, demonstrating that RynnWorld-Teleop serves as a high-fidelity, scalable data engine for the next generation of robotic agents.

The framework is built upon the Wan-I2V architecture (Wang et al., 2025a), which utilizes a 3D Variational Autoencoder (VAE) and a Transformer-based denoiser FΘ. The model is trained to predict the velocity field, conditioned on the reference-image latent zref = E(Iref) and a depth-aware skeletal control latent c:

LCFM = Et,z0,ϵ h

∥vΘ(zt, t, zref, c) − (ϵ − z0)∥

2

Depth-aware action representation is achieved by rendering 21-joint hand poses with camera-distance-modulated color and radius to resolve depth ambiguity. This rendered pose video is projected into the latent space using a pretrained VAE encoder to yield a control latent c ∈ R C×T ×H×W that is spatially and temporally aligned with the target video latent, facilitating fine-grained grounding between action and appearance.

The action-conditioned video generation utilizes an additive patch-embedding scheme with distribution alignment:

x = PatchEmbedzC→D(zt) + α · PatchEmbedcC→D(ec), ec = c − µcσc · σz + µz, (3)

where zt ∈ R C×T ×H×W is the noisy video latent and ec is the aligned control latent.

The progressive cross-domain training involves two stages: Stage 1 (Egocentric Human Pretraining) on datasets like EgoDex and VITRA to learn fundamental dynamics of hand-object interactions, followed by Stage 2 (Robotic Domain Adaptation) where the model is fine-tuned on paired teleoperation data, mapping human gestures to robotic actions via Inverse Kinematics (IK).

Autoregressive distillation converts the bidirectional teacher model into a causal student model capable of real-time interactive generation. This involves a Causal Flow-Matching Warm-up phase to establish temporal causality and a Distribution Matching Distillation (DMD) stage, which utilizes a learned critic and the frozen teacher to push the student’s output distribution toward that of the teacher, ensuring high-fidelity synthesis in only four denoising steps.

The system is instantiated as a complete digital teleoperation system where raw operator motion is converted into paired (video, action) trajectories through Retargeting—which computes target end-effector poses via a calibrated coordinate-transform chain and solves Inverse Kinematics (IK) using an iterative damped least-squares (DLS) solver, regularized by a null-space shoulder prior. Skeletal-Conditioned Synthesis renders the corresponding robotic execution video by conditioning RynnWorld-Teleop on the reference image and skeletal sequence. Mitigating drift is achieved through Chunked Re-anchoring, where generation is re-anchored by providing the actual egocentric frame from the robot’s camera at each subsequent chunk.

In evaluation, RynnWorld-Teleop was tested on four complex tasks: Dual Picking, Block Pushing, Bimanual Lifting, and Lid Placement. Results show that policies trained purely on data generated by RynnWorld-Teleop can transfer zero-shot to real robots with success rates up to 100.0% for Block Pushing and Bimanual Lifting. Furthermore, augmenting real demonstrations with digitally teleoperated data consistently raises success rates, and feature distribution analysis using t-SNE shows significant overlap between RynnWorld-Teleop-generated data and real-world trajectories, indicating the model successfully captures the underlying distribution of robotic manipulation. The system achieves an effective interactive frequency of ∼40 Hz on a single NVIDIA H100 GPU. The work demonstrates that digital teleoperation can serve as a high-fidelity data engine that both substitutes for and amplifies physical teleoperation in robot learning.

Improvements for AI systems

Based on the scientific paper RynnWorld-Teleop: An Action-Conditioned World Model for Digital Teleoperation, here are the specific improvements suggested for AI systems and what those improved systems can achieve:


  1. Improve Robot Data Collection Scalability via Digital Teleoperation Paradigm:

  2. Enable Zero-Shot Sim2Real Transfer Across Diverse Dexterous Tasks:

  3. Achieve High-Fidelity, Real-Time Interactive Video Synthesis for Control Loops:

  4. Implement Robust, Grounded Action Conditioning for Bimanual Coordination:

  5. The system can bypass the bottleneck of physical hardware and manual environment resets by replacing real robot demonstrations with an action-conditioned generative world model (RynnWorld-Teleop). This allows data collection to be driven purely by operator imagination, decoupling it from physical infrastructure constraints.

  6. The resulting robot policies trained exclusively on this synthesized data can achieve effective zero-shot Sim2Real transfer across a wide variety of dexterous and complex bimanual manipulation tasks (e.g., dual picking, block pushing). This means robots can learn complex skills directly from human intent without needing extensive, task-specific physical training.

  7. The system provides a real-time (40+ FPS) pipeline for synthesizing high-fidelity, egocentric videos conditioned on an operator's hand pose stream. This enables closed-loop interaction where the operator provides control signals that immediately drive the generation of robot trajectories, allowing for fluid, responsive digital teleoperation and real-time skill stringing.

  8. The system utilizes a depth-aware skeletal representation and distribution alignment techniques to ensure that every generated frame is precisely grounded in joint-level action signals, effectively bridging the gap between human intent (hand pose) and robotic execution (joint configuration). This allows the AI to maintain strict adherence to complex, high-precision bimanual coordination even when generalizing to unseen objects or backgrounds.

Sources

Related papers