World-Ego Modeling for Embodied Video Generation in Long-Horizon Navigation-Manipulation Tasks

arXiv:2605.19957 · cs.CV, cs.AI, cs.RO · Submitted 2026-05-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "World-Ego Modeling for Embodied Video Generation in Long-Horizon Navigation-Manipulation Tasks".

Jane: The gist The World-Ego Modeling paradigm decomposes embodied video prediction into persistent world regularities and robot-centric ego dynamics, which addresses long-horizon degradation in hybrid navigation-manipulation tasks.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we started by looking at how standard world models usually fail when you combine navigation and manipulation in a single video prediction stream one <ref:2605.19957#pg1>. They claim that this is because they predict everything as one thing, which leads to degradation over long stretches of time in hybrid tasks one <ref:2605.19957#pg1>.

Jane: The main thesis of the World-Ego Modeling paper is that you should decompose that prediction into two distinct components: the world evolution and the ego dynamics one <ref:2605.19957#pg1>. They argue that this helps because the world handles persistent scene regularities, while the ego handles things like robot behavior and instruction-conditioned actions one <ref:2605.19957#pg1>.

Lu: What matters is that this decomposition allows for a more interpretable way to understand what’s happening in a video sequence one <ref:2605.19957#pg1>. It lets us see which part of the prediction is due to general scene rules and which part is due to what the robot is specifically doing with its commands one <ref:2605.19957#pg1>.

Meng: So, if we look at an action sequence, we can isolate whether the model is predicting how the environment stays consistent or how the robot decides where to move next based on its instructions one <ref:2605.19957#pg1>. That seems like a clearer way to diagnose model weaknesses.

Lalam: For us as models, this means we can have specialized components for different kinds of information: one for scene patterns and one for action logic one <ref:2605.19957#pg1>. It’s about making the AI's internal representation more organized.

Tom: And they claim that by defining clear boundaries between these two parts—whether it's motion based or semantic based—we can get stable long-horizon rollouts for those hard hybrid tasks one <ref:2605.19957#pg1,stable long-horizon rollouts for>.

Jane: The paper sets up this framework by proposing three ways to draw that boundary: motion, semantic, and intention views one <ref:2605.19957#pg1>. They even adopt the semantic-based view as the default way to define the world and ego in their World-Ego Model one <ref:2605.19957#pg1>.

Lu: The motivation is that this decomposition addresses why standard models struggle with long-horizon scenarios when they have interleaved navigation and manipulation behaviors one <ref:2605.19957#pg1>. It’s about capturing those distinct dynamics separately.

Meng: So, the paper is essentially saying that instead of one monolithic prediction, we should train separate systems for world and ego so they don't interfere with each other over time one <ref:2605.19957#pg1>. That’s a practical engineering goal.

Lalam: It suggests a path forward where we can engineer better control by having explicit components for environment consistency and robot-centric dynamics one <ref:2605.19957#pg1>.

Tom: And that leads right into how they actually build the World-Ego Model, which is called WEM, coupling an implicit planner with a cascade parallel mixture of experts diffusion generator one <ref:2605.19957#pg1>.

Jane: They describe the prediction stage where a vision-language state predictor creates separate ego and world states for conditioning the main generator one <ref:2605.19957#pg1>. This predictor uses asymmetric query budgets and role-conditioned attention to keep those states focused on their specific predictive roles one <ref:2605.19957#pg1>.

Lu: Those design choices are key because they restrict the attention horizon for each group, anchoring the world state to accumulated scene regularities and the ego state to local instruction-driven dynamics one <ref:2605.19957#pg1>.

Meng: So, we're essentially using that predictor to feed context into a diffusion generator so it can generate a sequence that respects both those separated constraints one <ref:2605.19957#pg1>.

Lalam: It’s about conditioning the generation process on these two distinct conceptual views of the future evolution, which is an interesting way to structure the generation pipeline one <ref:2605.19957#pg1>.

Conclusion: Tom: So looking at the whole paper, it’s about solving that long-horizon problem by explicitly separating the world dynamics from the robot's ego dynamics one <ref:2605.19957#pg1>. The authors are Zuyao Lin and his team one <ref:2605.19957#pg1>.

Jane: They used this World-Ego Modeling approach, or WEM, to show a way to get stable predictions for hybrid navigation and manipulation tasks on their HTEWorld dataset one <ref:2605.19957#pg1>. It’s about achieving better consistency over longer time steps one <ref:2605.19957#pg1>.

Lu: The implication is that for embodied AI, we can move away from monolithic world models toward structured models where the environment dynamics are cleanly separated from the agent's specific control logic one <ref:2605.19957#pg1>.

Meng: Practically speaking, this means if we want a robot to perform a long sequence of actions involving both driving and grabbing objects, we can build a system that handles those two things more reliably one <ref:2605.19957#pg1>.

Lalam: It suggests that explicit world-ego separation is a promising direction for building systems that are structured and controllable in the embodied space one <ref:2605.19957#pg1>.

Tom: Exactly. The paper shows how this specific framing makes the model perform well on complex tasks, even when compared to models focused just on manipulation one <ref:2605.19957#pg1>.

Jane: It's about making sure the AI doesn't get lost in the details of one task while ignoring the constraints of the other one <ref:2605.19957#pg1>.

Lu: This structure opens up possibilities for future work where we can build on this by combining it with hierarchical planning or explicit memory refresh one <ref:2605.19957#pg1>.

Meng: I think that's where things could go next, because disentanglement alone might not be enough to solve every long-horizon problem, so we need to figure out how to integrate this structure further one <ref:2605.19957#pg1>.

Lalam: Maybe future work should focus on using weakly supervised methods or self-supervised ways to estimate those world and ego structures directly from video and interaction signals one <ref:2605.19957#pg1>.

Tom: So that’s the direction for the next phase: moving from demonstrating the concept to actually making these structured models robust enough for real-world use one <ref:2605.19957#pg1>.

Jane: That separation seems like a very promising path forward for creating more reliable embodied AI systems one <ref:2605.19957#pg1>.

Institute of Automation, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Zhongguancun Academy 4Shanghai Jiaotong University 5Peking University

cs.CV, cs.AI, cs.RO

Submitted: 2026-05-19

Updated: 2026-10-08

Code: https://github.com/ZGCA-HMI-Lab/WEMhttps:

Project page: https://zgca-hmi-lab.github.io/WEMhttps://github.com/ZGCA-HMI-Lab/WEMhttps://huggingface.co/Zoorao/WEMhttps://huggingface.co/datasets/Zoorao/HTEWorld

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 91/100

The gist: The gist The World-Ego Modeling paradigm decomposes embodied video prediction into persistent world regularities and robot-centric ego dynamics, which addresses long-horizon degradation in hybrid

Key concepts

World and Ego Decomposition
This paradigm splits future video prediction into two parts: the world, which captures persistent scene regularities independent of instructions, and the ego, which represents robot-centric dynamics conditioned by current instructions. This separation helps maintain coherence over long sequences where traditional models struggle.
World-Ego Boundary Views
The paper defines how to draw this split using three methods: motion-based (separating scene flow from contact dynamics), semantic (defining the ego as the robot and manipulated object), and intention-based (distinguishing history from current instruction). The semantic view is chosen as the default for WEM.
CP-MoE Generator
The generation stage uses a cascade-parallel mixture-of-experts diffusion model. This architecture splits the main model into a shared expert and specialized rear stages for the world and ego. This allows the generator to predict a world-ego proxy by conditioning it on both separate state predictions.
Full Disentanglement
This is a strategy used in WEM's generation stage where tokens are processed through routing, expert specialization, and unrouting. It aims for the strongest structural separation between the world and ego components during video generation, which was found to yield the best results.

Terminology

Summary

The gist The World-Ego Modeling paradigm decomposes embodied video prediction into persistent world regularities and robot-centric ego dynamics, which addresses long-horizon degradation in hybrid navigation-manipulation tasks.

World-Ego Modeling Paradigm

World models typically predict distinct evolutions of the world and the ego within a single stream, leading to degradation in long-horizon embodied scenarios, particularly in hybrid tasks with interleaved navigation and manipulation behaviors<ref:2605.19957#pg3>. World-Ego Modeling introduces a new conceptual paradigm that decomposes future evolution into world and ego components<ref:2605.19957#pg2>. This decomposition is motivated by the fact that the world captures persistent instruction-agnostic scene regularities while the ego captures robot-centric, instruction-conditioned dynamics<ref:2605.19957#pg3>. The paper defines this boundary from three perspectives: i) motion-, semantic-, and intention-based views<ref:2605.19957#pg3>.

Defining World and Ego Boundaries

The three operational definitions for drawing the world-ego boundary are as follows<ref:2605.19957#pg5>:

  1. Motion-based view draws the boundary by the source of visual motion, where pixels whose motion matches this scene flow are explained by viewpoint change alone and are assigned to the world, while those reflecting contact-driven object dynamics induced by the embodiment and assigned to the ego<ref:2605.19957#pg5>.

  2. Semantic-based view draws the boundary by the embodied role of scene entities, where the robot itself and any object currently being manipulated jointly constitute the ego region, capturing robot-centric, instruction-conditioned dynamics<ref:2605.19957#pg5>.

  3. Intention-based view draws the boundary at the source of conditioning information, where world reflects what is established by visual history and ego reflects what is induced by the current instruction<ref:2605.19957#pg5>. The paper adopts the semantic-based view as the default world-ego definition in WEM<ref:2605.19957#pg5>.

World-Ego Model (WEM) Architecture

The World-Ego Model (WEM) instantiates this paradigm as a unified embodied world model that couples an implicit separate world-ego planner with a cascade-parallel mixture-of-experts (CP-MoE) diffusion generator<ref:2605.19957#pg2>. The framework consists of two stages: a prediction stage that infers separate ego and world states via the vision-language state predictor, and a generation stage based on the CP-MoE generator<ref:2605.19957#pg4>.

The prediction stage uses a vision-language state predictor that augments a pretrained VLM with ego/world queries to produce S e k and S w k, which serve as conditioning signals for the generator<ref:2605.19957#pg8>. This state predictor employs two design choices for world-ego separation at the state level: asymmetric query budgets and role-conditioned attention (RCA)<ref:2605.19957#pg8>. RCA restricts the attention horizon of each group to its predictive role, anchoring S w k to scene regularities accumulated from history and S e k to instruction-conditioned dynamics in the current local context<ref:2605.19957#pg8>.

Generation Stage and Disentanglement Strategies

The generation stage uses a CP-MoE generator that splits the DiT backbone into a shared preceding expert and specialized rear stages (ego and world experts)<ref:2605.19957#pg5>. The preceding expert is conditioned on both states to predict a world-ego proxy, which operationalizes the world-ego boundary defined in Section 3.2<ref:2605.19957#pg5>.

The paper explores three disentanglement strategies for the generation stage:

  1. Pre-disentanglement routes tokens before the rear stage, where a single rear module with restricted cross-attention is used<ref:2605.19957#pg6>.

  2. Post-disentanglement fuses the outputs of separate ego and world experts, where each branch processes the full token sequence under its respective state<ref:2605.19957#pg6>.

  3. Full disentanglement combines routing, expert specialization, and unrouting for stronger structural separation<ref:2605.19957#pg6>. The paper adopts full disentanglement as the default instantiation of WEM<ref:2605.19957#pg10>.

Evaluation and Results

To enable rigorous evaluation, the paper constructs HTEWorld, the first benchmark for long-horizon world modeling with hybrid navigation-manipulation tasks, providing 125K video clips (over 4.5M frames) with fine-grained action annotations and 300 multi-turn evaluation trajectories<ref:2605.19957#pg2>. WEM achieves state-of-the-art performance on HTEWorld while remaining competitive on existing manipulation-only benchmarks<ref:2605.19957#pg2>.

Design Study I showed that the semantic-based view achieves the best EWMScore, outperforming the motion-based and intention-based views by 2.12 and 2.79 points, respectively<ref:2605.19957#pg9>. Design Study II demonstrated that full disentanglement performs best, yielding an EWMScore of 61.48<ref:2605.19957#pg10>. WEM outperforms representative baselines on HTEWorld while remaining competitive on manipulation-only benchmarks<ref:2605.19957#pg11>.

Limitations and Future Work

WEM is only an initial instantiation of the broader World-Ego Modeling paradigm, and the current study explores a limited set of boundary definitions, disentanglement designs, and simulated embodied settings<ref:2605.19957#pg12>. Limitations include the lack of real-world evaluation to assess sim-to-real generalization<ref:2605.19957#pg22> and that world-ego disentanglement alone is not sufficient to fully solve long-horizon embodied generation, suggesting future work should combine WEM with hierarchical planning or explicit memory refresh. The paper suggests that future work should develop stronger weakly supervised or self-supervised boundary estimation methods to infer world-ego structure directly from video, language, and interaction signals<ref:2605.19957#pg22>.

The paper concludes that World-Ego Modeling is a promising direction for structured and controllable embodied world models<ref:2605.19957#pg12>. The final sentence of the conclusion states that WEM outperforms representative baselines on HTEWorld while remaining competitive on manipulation-only benchmarks, suggesting explicit world-ego separation as a promising direction for structured and controllable embodied world models.

Improvements for AI systems

  1. World-Ego Modeling (WEM) can be instantiated to perform long-horizon multi-turn embodied evolution in hybrid navigation-manipulation tasks, enabling agents to maintain scene consistency during long sequences of interleaved navigation and manipulation behaviors.

  2. The system can achieve improved scene geometry preservation and instruction alignment by adopting the semantic-based view with full disentanglement, as this configuration was shown to yield the highest EWMScore (61.48) compared to other boundary definitions.

  3. The agent can exhibit stronger phase-matched motion and long-horizon stability by leveraging the HTEWorld-specific metrics like PMPA, CPDM, and FPHS, which measure better phase-matched motion, camera/object coordination, and long-horizon stability.

  4. A new annotation pipeline can be utilized to generate distinct, context-rich prompts for language supervision by fusing robot grounding with episode trajectory context and temporal phase hints like clip’s index within its parent action (e.g., 'clip 3 of 5, 60% through this action'), leading to more accurate visual descriptions.

  5. The system can be adapted for downstream tasks like autonomous driving or robot planning by separating persistent environmental structure from self-conditioned dynamics, which allows for more effective modeling of road-scene regularities from ego-vehicle behavior.

Abstract

Embodied video world models typically capture both scene evolution and the robot's behavior, which we refer to as the world and the ego, respectively. The world and the ego exhibit different underlying dynamics: world prediction relies primarily on visual history and emphasizes scene stability, whereas ego prediction relies more strongly on the current instruction and emphasizes accurate instruction following. Modeling both components within a single generation stream can entangle these different dependencies, making it difficult to specialize the prediction of either component. Consequently, it becomes difficult to simultaneously maintain scene consistency and accurate instruction following, particularly in long-horizon navigation-manipulation tasks. In this paper, we propose to decompose an embodied video into the world and the ego and disentangle their generation processes. Specifically, we define the world as the background and currently unmanipulated objects, and the ego as the robot and currently manipulated objects. Based on this definition, we develop the World-Ego Model (WEM), which combines a vision-language state predictor using role-conditioned attention (RCA) and asymmetric query budgets with a semantic-routed mixture-of-experts (SR-MoE) diffusion generator. To enable rigorous evaluation, we further construct HTEWorld, a dataset and benchmark for long-horizon embodied video generation with hybrid navigation-manipulation tasks, providing 125K training video clips comprising over 4.5M frames with fine-grained instructions, together with 300 multi-turn evaluation trajectories covering over 2K instructions. Extensive experiments show that WEM achieves state-of-the-art performance on HTEWorld while remaining competitive on existing manipulation-oriented evaluations.

Sources

Related papers