AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
AlayaWorld Team, Kaipeng Zhang, Chuanhao Li, Yifan Zhan, Yongtao Ge, Yuanyang Yin, Jiaming Tan, Kang He, Liaoyuan Fan, Mingliang Zhai, Ruicong Liu, Xiaojie Xu, Xuangeng Chu, Zhen Li, Zhengyuan Lin, Zhixiang Wang, Zian Meng, Zihui Gao
Alaya Lab
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: Authors are listed alphabetically by the first name and their role. See the contribution section for details
Code: https://github.com/AlayaLab/AlayaWorld
Project page: https://alaya-lab.github.io/AlayaWorld
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) presents an improved version of AlayaWorld.
Terminology
Summary
AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1) presents an improved version of AlayaWorld. While the backbone architecture, chunkwise autoregressive generation scheme, and training data remain unchanged from the previous release, the report substantially revises how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure.
To this end, two major changes are made. First, the previous depth-warping-based spatial memory is replaced with a streaming 3D point-cache renderer. Second, the conditioning pipeline is redesigned so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.
The report details the six changes as follows:
-
Motion-aware image conditioning: "In the previous version, an image condition was encoded as an isolated frame, yielding a latent representation without the temporal context present in causally encoded video latents. We instead construct a nine-frame window consisting of the conditioning frame and its preceding eight frames (stride + 1 = 9), encode the full window with the causal VAE, and use the second latent as the image condition. The single conditioning frame is retained only as a decoder prefix. The same nine-frame encoding pattern is used for chunk-to-chunk handoff during inference, so image conditioning and autoregressive continuation are represented with matched temporal context."
-
Causal encoding of a streaming 3D spatial memory: "We replace the previous DA3-based depth warping [1] with a streaming 3D point cache, using ViGeo [3] to estimate the per-pixel 3D geometry. After each generated chunk, per-pixel 3D points are registered into a persistent cache, which is then re-rendered from the planned viewpoint of the next chunk. To maintain geometric consistency over the rollout, the camera trajectory is scale-aligned to the cache using a pairwise-median ratio over prefix displacements, while camera intrinsics are estimated once per clip by least-squares fitting of a pinhole camera model to the point map and then kept fixed. We also change how the rendered spatial condition is encoded. Instead of encoding rendered frames independently, we prepend a real prefix frame and encode the resulting sequence causally. The prefix latent is then discarded, leaving spatial-memory latents that follow the same causal encoding structure as the target video latents. Tokens corresponding to invalid rendered regions are removed rather than masked, preserving compatibility with FlashAttention."
-
Pixel-aligned temporal memory: "Temporal memory is reduced from six latents to four and is constructed by aligning the memory window in pixel space before VAE encoding. For N = 4 memory latents, we encode exactly 1 + (N − 1) × 8 = 25 frames, producing exactly four causal latents. The data loader and trainer use the same frame-to-latent convention to ensure consistent window boundaries. This removes the off-by-one boundary latents that could previously leak into the temporal-memory context."
-
Hard memory dropout: "The previous memory-dropout strategy zeroed memory tokens while preserving their positions in the sequence. Although their values were removed, the sequence structure still revealed whether memory tokens were present. We replace this with hard dropout, which removes the memory tokens entirely and therefore changes the actual sequence length. Correspondingly, the first rollout step at inference is performed without temporal memory, matching the memory-free cases seen during training."
-
A unified VAE protocol across training and inference: "We standardize the pixel–latent interface used throughout training, autoregressive rollout, and evaluation. Chunk handoff consistently uses an anchor frame together with the same nine-frame causal encoding window. Decoding supports both direct latent continuation and an RGB decode–re-encode path, while evaluation supplies ground-truth continuations together with the two causal prefix latents required by the VAE. This removes several stage-specific encoding and decoding conventions from the previous version and reduces discrepancies between training and autoregressive inference."
-
Geometry-based camera control: "The dedicated camera AdaLN branch is removed. Instead, the planned camera trajectory is used to determine the viewpoint from which the 3D point cache is re-rendered, and the resulting spatial condition provides the camera guidance to the model. After causal-VAE encoding, viewpoint information is therefore presented through the same visual latent representation as other spatial conditions. This directly couples camera control to scene geometry, including scale, visibility, and parallax, without requiring a separate camera-conditioning side channel. Text conditioning continues to specify scene semantics, while the rendered spatial memory specifies the desired viewpoint."
The experimental results are reported on the WBench [2] navigation split, covering 158 navigation cases across Video Quality, Setting, Interaction, Consistency, and Physical metrics. AlayaWorld performs strongly on the Consistency benchmarks, achieving the best overall Consistency score of 89.5. It also remains competitive in Video Quality, with an average score of 79.1, while maintaining solid navigation performance under the interactive evaluation protocol. Overall, the results indicate that AlayaWorld is effective at preserving visual and geometric consistency during long-horizon interactive generation.
For Video Quality, AlayaWorld achieves the best Imaging score of 67.7 and remains competitive in Aesthetic quality with a score of 62.6. Its overall Video Quality score of 79.1 is close to the strongest competing methods, despite not leading on temporal metrics such as Flickering, Dynamic, and Smoothness. These results suggest that AlayaWorld maintains high perceptual quality and image fidelity during autoregressive interaction, while leaving room for further improvement in short-term temporal dynamics.
The main advantage of AlayaWorld is observed in Consistency. It achieves the best results in Background Consistency, Perspective Consistency, Subject Consistency, and Geometric Consistency, and ranks second in both Spatial and Gated Spatial Consistency. In particular, the strong Perspective and Geometric scores indicate that AlayaWorld better preserves scene structure and viewpoint-dependent geometry as the camera moves through the environment, while the high Background and Subject scores show that previously observed visual content remains stable over extended interactions. Together, these improvements lead to the highest overall Consistency score among all evaluated methods, validating the effectiveness of the proposed spatial and temporal memory mechanisms for long-horizon generation.
For Interaction, AlayaWorld achieves a Navigation score of 79.9, demonstrating responsiveness to navigation controls, although competing methods perform better. Its performance on the Setting and Physical categories is comparatively weaker, particularly for Scene consistency and Causal Fidelity. These results suggest that while AlayaWorld provides strong visual persistence and geometric stability, improving environment-level semantic preservation and physical interaction modeling remains an important future direction.
Qualitative results on WBench show that across diverse navigation scenarios, AlayaWorld produces coherent visual transitions while maintaining the appearance of scene content and the spatial relationships among objects throughout the interaction. The results further show stable scene structure under continuous viewpoint changes, demonstrating strong visual persistence and spatial consistency during interactive navigation. Additional qualitative results across a diverse set of scenes demonstrate that the model generalizes well across different environments and maintains consistent visual content under varying scene layouts and motion patterns.
Improvements for AI systems
Improvements to AI Systems:
-
Latent-Space Temporal Alignment for Conditioning: Encode conditioning inputs (e.g., single images, spatial memories) using the same causal-VAE with matched temporal windows (e.g., 9-frame stride) as the generated video. This ensures conditioning latents carry identical temporal statistics to target latents, reducing train-inference mismatch and improving autoregressive consistency.
-
Streaming 3D Point-Cache Memory with Geometry-Aware Rendering: Replace depth-warping or static-frame memory with a persistent 3D point cloud that is incrementally updated per generated chunk and re-rendered from the next planned viewpoint. This enables long-horizon spatial consistency, scale alignment via pairwise-median ratios, and fixed intrinsics estimation, allowing the model to maintain coherent geometry across arbitrary camera paths.
-
Hard Memory Dropout for Robust Rollout: During training, randomly remove entire memory tokens (not just zero them) to alter sequence length. At inference, the first rollout step uses no temporal memory, matching training distribution. This improves generalization to memory-free starts and prevents the model from relying on positional hints of memory presence.
-
Unified VAE Encoding/Decoding Protocol: Standardize pixel–latent interfaces across training, rollout, and evaluation (e.g., always use anchor frame + 9-frame causal encoding for handoff; support both direct latent continuation and RGB decode–re-encode). This eliminates stage-specific discrepancies, reducing error accumulation during long autoregressive generation.
-
Geometry-Based Camera Control without Side Channels: Remove separate camera-conditioning branches (e.g., AdaLN) and instead derive viewpoint control entirely from re-rendered 3D point-cache images, encoded into the same visual latent space. This couples camera motion to scene geometry (scale, visibility, parallax), improving perspective and geometric consistency without extra parameters.
-
Pixel-Aligned Temporal Memory Windows: Construct temporal memory by encoding exactly 1 + (N−1)×8 frames (e.g., 25 frames for N=4 latents) in pixel space before VAE encoding, ensuring frame-to-latent boundaries are identical across data loading and training. This removes off-by-one boundary latents that leak incorrect temporal context.
What the Improved AI System Can Do:
-
Long-Horizon Interactive Video Generation: Generate coherent, geometrically stable video rollouts over hundreds of frames with consistent scene structure, background, and subject appearance under continuous user-controlled camera navigation.
-
Accurate Viewpoint-Dependent Rendering: Maintain correct perspective, parallax, and object occlusion as the camera moves, using a persistent 3D point cache that is re-rendered per step, without needing explicit depth maps or separate camera embeddings.
-
Seamless Autoregressive Continuation: Produce smooth transitions between chunks with no temporal-context mismatch, thanks to matched causal-VAE encoding windows and unified handoff protocols.
-
Robust Memory-Free Initialization: Start generation from a single image or text prompt without temporal memory, while still achieving high consistency after the first step, due to hard memory dropout training.
-
High Consistency in Interactive Environments: Achieve top-tier scores in background, perspective, subject, and geometric consistency during long navigation tasks, outperforming prior methods in preserving scene identity and spatial relationships.
-
Efficient Memory Usage: Remove invalid rendered regions (e.g., out-of-view points) rather than masking them, enabling compatibility with FlashAttention and faster inference on long sequences.
Sources
- Depth Anything 3: Recovering the Visual Space from Any Views
- WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation
- Towards Consistent Video Geometry Estimation
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection