WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models

arXiv:2605.25077 · cs.CV · Submitted 2026-05-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "WorldCraft: From Camera Navigation to Object Manipulation in Interactive Video World Models".

Jane: The gist The WorldCraft framework expands interactive video world models from camera navigation to object-level trajectory actions, enabling users to manipulate selected objects while continuing camera navigation.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: To get into the specifics, WorldCraft proposes generating future frames where a selected object follows a user-drawn path while the camera keeps moving. It's all about that composable control they call it.

Jane: They achieved this by introducing three main pieces: Normalized World Trajectory, Spatial-Pathway LoRA, and Trajectory-Anchored State Persistence. That’s the core mechanism they put together to make it happen.

Lu: The concept of Normalized World Trajectory is really clever because it takes a trajectory drawn in user space and re-projects it based on where the camera is currently looking, which separates the object's movement from just the camera shifting around.

Meng: So, if I’m driving a virtual car and I want to steer left while simultaneously looking at the scenery change, this framework is designed to handle that specific kind of interaction.

Lalam: It sounds like they are essentially creating a system where you can draw a path for an object in the world, and the AI figures out how to make that object move along it while you keep looking around.

The paper's summary: Tom: The summary of WorldCraft is that they take existing models like WorldPlay and inject this object-level trajectory action capability into them without breaking the original camera control functionality.

Jane: What’s interesting is how they show that you can add object manipulation without overwriting the base camera controller, even though both things seem to use similar underlying parts of the model.

Lu: They analyzed that their analysis shows that the camera and trajectory control actually share the same spatial pathway inside the transformer structure, which is a key insight for how this works.

Meng: That means they aren't completely rebuilding the core navigation part; they are just adding a specific instruction pathway to it. That’s an important practical detail for implementation.

Lalam: The paper shows that you can adapt only the action encoder and projection layers to add trajectory control, which keeps the camera controller exactly as it was before.

The paper's improvements: Tom: They suggest a three-stage progressive training schedule to teach the model this new behavior gradually, starting with static cameras and moving up to dynamic sequences.

Jane: That progressive training is smart because it helps them avoid those messy interference issues where trying to learn object movement messes up the camera control, or vice versa.

Lu: They use a specific data curation pipeline that extracts video, camera, mask, and trajectory tuples from unlabeled video at scale. That structured supervision is what makes the learning possible.

Meng: Extracting those specific tuples automatically from raw video is a big hurdle for any practical application because you need accurate masks and ground truth trajectories to train it properly.

Lalam: They also mention using a scalable data curation pipeline involving VLM-guided discovery and multi-frame tracklet matching to get that structured supervision ready for the training process.

Conclusion: Tom: So, WorldCraft wraps up by showing they can support long-horizon autoregressive generation while keeping both scene consistency and object state preservation.

Jane: The big implication here is that these models aren't just good at watching a video; they can now be used for active manipulation in a way that respects the camera view.

Lu: They’ve shown that you can achieve long trajectory generation, like about ten point five seconds at twenty-four frames per second, with this composable control <ref:2605.25077#pg2>.

Meng: For someone building an application on top of this, it means you can create interactive simulations where the object moves predictably while the user is still navigating the environment.

Lalam: WorldCraft takes these models toward something that supports not just watching scenes but actively manipulating objects within them, which is really relevant for embodied AI development.

The Hong Kong University of Science and Technology AI Technology Center, Tencent Video

cs.CV

Submitted: 2026-05-24

Updated: 2026-10-08

Importance score: 89/100

The gist: The gist The WorldCraft framework expands interactive video world models from camera navigation to object-level trajectory actions, enabling users to manipulate selected objects while continuing

Key concepts

Normalized World Trajectory (NWT)
NWT takes user-drawn motions in a world coordinate system that doesn't change with the camera's view. It projects these desired movements onto the current camera view at every step. This separates the object's intended path from any distortion caused by how the camera is positioned, ensuring consistent control regardless of where you look.
Spatial-Pathway LoRA (SP-LoRA)
This mechanism adds object manipulation capabilities without breaking the existing camera controller. It applies low-rank updates only to specific layers—the action encoder and ProPE projection layers—which form the shared spatial control pathway. This allows the model to learn how to follow trajectories while keeping its fundamental ability to navigate scenes intact.
Trajectory-Anchored State Persistence (TASP)
TASP treats the object's path as a permanent spatial state in memory. When an object moves off-screen, TASP uses this trajectory information as a global 'where' signal. This signal refreshes the model's memory, allowing moved objects to accurately reappear at their new locations when they return to view.

Terminology

Summary

The gist The WorldCraft framework expands interactive video world models from camera navigation to object-level trajectory actions, enabling users to manipulate selected objects while continuing camera navigation.

WorldCraft's Core Contribution

WorldCraft introduces a framework that equips an interactive video world model with object-level trajectory actions while preserving its camera-control capabilities<ref:2605.25077#pg12>. It achieves this by generating future frames in which the selected object follows the prescribed trajectory as the camera simultaneously navigates the scene<ref:2605.25077#pg4>. The framework augments a backbone like WorldPlay through three novel components: Normalized World Trajectory (NWT), Spatial-Pathway LoRA (SP-LoRA), and Trajectory-Anchored State Persistence (TASP)<ref:2605.25077#pg12>.

Key Mechanisms

The framework operates through a trajectory-centric control pipeline<ref:2605.25077#pg4>. The three main components are defined as follows:

  1. Normalized World Trajectory (NWT): This component represents user-drawn motion in a camera-invariant world coordinate system and dynamically re-projects it under the current camera pose, separating object motion from camera-induced screen-space displacement<ref:2605.25077#pg5>. It lifts user trajectories into a normalized coordinate on the first frame and then re-projects them under the current camera pose at each generation step<ref:2605.25077#pg5>.

  2. Spatial-Pathway LoRA (SP-LoRA): This mechanism inject[s] this world-space signal through the model’s spatialcontrol pathway, adding object manipulation capability while preserving the pretrained camera controller<ref:2605.25077#pg12>. It adapts only the action encoder and ProPE projection layers to add trajectory control without disrupting the base camera controller<ref:2605.25077#pg5>.

  3. Trajectory-Anchored State Persistence (TASP): This component "treat[s] the world trajectory as a persistent spatial state and refreshes autoregressive memory after trajectory-conditioned generation, allowing moved objects to reappear at their updated positions after leaving the camera view"<ref:2605.25077#pg12>. TASP uses the world-space trajectory as a global “where” signal that complements autoregressive memory’s “what” signal when the object leaves and re-enters view<ref:2605.25077#pg5>.

Technical Details and Training

WorldCraft builds on WorldPlay [23], an autoregressive video world model based on HunyuanVideo1.5 [16]<ref:2605.25077#pg4>. The trajectory injection is done by replacing the image-conditioning channels with first-frame latent features displaced to the target positions, creating a condition cˆtraj<ref:2605.25077#pg5>. Training is conducted using a three-stage progressive schedule<ref:2605.25077#pg5>. Stage 1 trains trajectory control on static-camera data using BI attention and SP-LoRA<ref:2605.25077#pg5>. Stage 2 extends training to dynamic-camera sequences with AR attention<ref:2605.25077#pg5>. A scalable data curation pipeline is used to extract (video, camera, mask, trajectory) tuples from unlabeled video at scale<ref:2605.25077#pg5>.

Experimental Validation

Experiments show that WorldCraft enables accurate object control and preserves the video-based world model’s camera fidelity under camera-only evaluation<ref:2605.25077#pg12>. Quantitative results on trajectory control under static camera demonstrate that WorldCraft achieves the lowest trajectory error (TA) while simultaneously producing the best pixel fidelity and semantic consistency across all test clips<ref:2605.25077#pg5>. Furthermore, on camera-only inputs, WorldCraft retains the camera-control capability of the base model, with its RPErot being 0.131 at 61 frames compared to 0.120 for WorldPlay<ref:2605.25077#pg5>. Qualitative comparisons show that WorldCraft maintains scene consistency and, via TASP, recovers off-camera object state when the camera returns<ref:2605.25077#pg5>. The study concludes that WorldCraft supports long-horizon autoregressive generation with both scene-level consistency and object-level state preservation<ref:2605.25077#pg5>.

Limitations and Impact

The limitations include the fact that the mechanism only persists the state of entities with user-specified trajectories, leaving predicting uninstructed dynamics an open problem<ref:2605.25077#pg5>. Camera-trajectory compensation relies on monocular depth estimation, which introduces projection error at large camera rotations<ref:2605.25077#pg5>. The trajectory control operates at the granularity of latent tokens, limiting precision for very small objects<ref:2605.25077#pg5>. WorldCraft takes a step toward world models that support not only passive observation but active manipulation, a capability relevant to embodied AI, content creation, and simulation<ref:2605.25077#pg5>.

How it works

  1. NWT determines what coordinates to inject by lifting user trajectories into a normalized world-space coordinate system and dynamically re-projecting them under the current camera pose<ref:2605.25077#pg5>.

  2. SP-LoRA determines which parameters to adapt by applying low-rank updates only to the action encoder and ProPE projection layers, which are the shared spatial-control pathway<ref:2605.25077#pg5>.

  3. TASP determines how memory interacts with the injected signal across chunks by using the world-space trajectory as a persistent spatial state signal and employing pre-exit memory filtering<ref:2605.25077#pg5>.

A Progressive training

The three-stage pipeline systematically avoids both failure modes by gradually increasing data complexity and constraining the attention mode<ref:2605.25077#pg5>. Stage 1 trains trajectory control on static-camera data using BI attention and layer-selective LoRA<ref:2605.25077#pg5>. Stage 2 extends to dynamic-camera data with AR attention<ref:2605.25077#pg5>. This strategy ensures that the camera effect direction is highly stable regardless of trajectory input, confirming that pathway-selective LoRA achieves an asymmetric decoupling where camera control is preserved while trajectory control is added<ref:2605.25077#pg5>.

A Scalable data curation pipeline

The pipeline adapts to two data sources with complementary strengths: VLM-guided discovery (WISA-80K) and multi-frame tracklet matching (SpatialVID-HQ)<ref:2605.25077#pg5>. This involves camera estimation via ViPE, subject discovery using VLM and GroundingDINO, and trajectory extraction using CoTracker3<ref:2605.25077#pg5>. The training data is standardized to 30 fps and 97 frames (≈3.2 s)<ref:2605.25077#pg5>.

A Summary of Contributions

The contributions are threefold<ref:2605.25077#pg12>:

  1. Object-level actions for interactive video world models, formulated as a new action modality for autoregressive video world models<ref:2605.25077#pg5>.

  2. Normalized World Trajectory, which lifts user trajectories into a normalized world-space coordinate system and dynamically re-projects them under the current camera pose<ref:2605.25077#pg5>.

  3. Off-camera state prediction, which uses the world-space trajectory as a persistent spatial state signal for off-camera objects and refreshes autoregressive memory so moved objects reappear at their updated positions<ref:2605.25077#pg5>.

Conclusion

WorldCraft introduces a framework that extends camera-controlled video world models with precise object-level action control<ref:2605.25077#pg12>. It identifies the shared spatial-control pathway underlying camera motion and object trajectories, and adapts it with a lightweight pathway-selective LoRA to add trajectory controllability while preserving the base model’s camera fidelity<ref:2605.25077#pg5>. Together with TASP-based memory refresh and progressive training, WorldCraft supports long-horizon autoregressive generation with both scene-level consistency and object-level state preservation<ref:2605.25077#pg5>. This points toward interactive world models that can not only navigate scenes, but also manipulate and reason about objects within them<ref:2605.25077#pg5>.

References

[1] GameGen-X Authors. Gamegen-x: Interactive open-world game video generation. arXiv preprint arXiv:2411.00769, 2024<ref:2605.25077#pg12>.

[3] David Ha and Jürgen Schmidhuber.

Improvements for AI systems

  1. Bold Header: Object-level action modality introduction

The framework WorldCraft introduces object-level trajectory actions as a new action modality for autoregressive video world models, enabling users to manipulate selected entities while continuing camera navigation. This allows users to achieve composable camera-object control.

  1. Bold Header: Camera invariance via normalized world trajectory

Normalized World Trajectory (NWT) lift[s] user trajectories into a normalized coordinate on the first-frame reference plane and re-projects it under the current camera pose, which is described as yielding a camerainvariant representation that disentangles ego-motion from object motion.

  1. Bold Header: Non-destructive parameter adaptation

The system utilizes Spatial-Pathway LoRA (SP-LoRA) to adapt only the spatial control pathway, ensuring camera and trajectory control share the same spatial pathway inside the transformer while preserving camera fidelity, as demonstrated by low relative weight changes in critical modules.

  1. Bold Header: Persistent off-camera state prediction

Trajectory-Anchored State Persistence (TASP) resolves off-camera state prediction by using the world-space trajectory as a persistent spatial signal for off-camera objects and refreshes autoregressive memory so moved objects reappear at their updated positions.

  1. Bold Header: Data curation pipeline for structured supervision

A Scalable data curation pipeline is established that automatically extracts video, camera, mask, and trajectory tuples from unlabeled video, which provides the structured supervision our method requires for training.

  1. Bold Header: Progressive training strategy

A three-stage progressive training schedule is employed to avoid interference between camera control and object trajectory control, including Stage 0 adaptation on real data without trajectory conditioning, Stage 1 for static cameras with SP-LoRA, and Stage 2 for dynamic camera sequences with AR attention.

  1. Bold Header: High fidelity under long-horizon rollouts

WorldCraft maintains scene consistency and, via TASP, recovers the goose at the correct off-camera-updated position when it re-enters view, demonstrating its ability to support long trajectory (∼10.5 s at 24 fps) with composable camera-object control.

  1. Bold Header: Robustness against camera rotation

The counterfactual analysis confirms that camera effect direction is highly stable regardless of trajectory input, as measured by a mean cosine similarity of 0.89, proving that the added trajectory perturbation does not alter the continuous camera pathway.

Sources

Related papers