Addressable Memory for Video World Models
Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taixé, Despoina Paschalidou, Jonathan Lorraine, Aljoša Ošep
cs.CV, cs.LG
Submitted: 2026-08-07
Updated: 2026-08-10
Comments: Project page: https://research.nvidia.com/labs/sil/projects/WorldTrace/
Code: https://github.com/NVIDIA/kvpress
Project page: https://research.nvidia.com/labs/sil/projects/WorldTrace
License: http://creativecommons.org/licenses/by/4.0/
The gist: We study visual persistence in interactive video world models.
Terminology
Abstract
We study visual persistence in interactive video world models. These models rely on a Key-Value (KV) cache as a growing visual memory to carry forward previously generated frames. However, we find that models can no longer reliably address stored content once rollouts extend beyond the training horizon, because temporal Rotary Positional Embeddings (RoPE) offsets then fall outside the range seen during training and the model struggles to retrieve the relevant visual information through attention. Moreover, naively compressing the cache in the RoPE-rotated space corrupts memory by averaging together incompatible positional phases. To address this, we propose WorldTrace, a training-free memory framework for long-horizon visual persistence. WorldTrace keeps compressed memory addressable by assigning each summary slot a distinct, in-distribution virtual position. Within this addressable cache, we study two memory compression approaches: WorldTrace-Field compresses history for temporal coherence, while WorldTrace-Landmark stores verbatim scene traces at detected transitions for episodic recall. We further introduce LoopBench, a benchmark evaluating whether a compressed cache can reconstruct a previously visited scene after a long detour. WorldTrace-Field improves temporal consistency by +15.5%, and WorldTrace-Landmark improves episodic recall by +19.5% on LoopBench, extending visually persistent generation without retraining.
Sources
- Cosmos 3: Omnimodal World Models for Physical AI
- Round and Round We Go! What makes Rotary Positional Encodings useful?
- NVIDIA OmniDreams: Real-Time Generative World Model for Closed-Loop Autonomous Vehicle Simulation
- Longformer: The Long-Document Transformer
- Variance Reduction for Expectations with Diffusion Teachers
- Mixture of Contexts for Long Video Generation
- Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion
- Grounded Forcing: Bridging Time-Independent Semantics and Proximal Dynamics in Autoregressive Video Synthesis
- Extending Context Window of Large Language Models via Positional Interpolation
- Context Forcing: Consistent Autoregressive Video Generation with Long Context
- VRAG: Learning World Models for Interactive Video Generation
- LoL: Longer than Longer, Scaling Video Generation to Hour
- Self-Forcing++: Towards Minute-Scale High-Quality Video Generation
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- MemCam: Memory-Augmented Camera Control for Consistent Video Generation
- Contextual Position Encoding: Learning to Count What's Important
- Long-Context Autoregressive Video Modeling with Next-Frame Prediction
- Memorize When Needed: Decoupled Memory Control for Spatially Consistent Long-Horizon Video Generation
- A$^2$ATS: Retrieval-Based KV Cache Reduction via Windowed Rotary Position Embedding and Query-Aware Vector Quantization
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models