LOCI: Spatial Linear Memory for Streaming World Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LOCI: Spatial Linear Memory for Streaming World Models".
Jane: When a camera revisits a previously observed region, a video world model must reproduce what was there before,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're looking at the paper titled "LOCI: Spatial Linear Memory for Streaming World Models," which sounds like it tackles a really tough problem in making video models remember things correctly when they revisit a scene. Jane, can you explain what that title actually means in plain English?
Jane: Well, essentially the title points to how this system uses spatial memory—meaning it stores information based on where things are in space—to handle streaming world models that need to recall what they saw before. It’s focused on maintaining that spatial context across long videos so the model doesn't lose sight of where it is going.
Lu: I think what's interesting about this title is the emphasis on "Spatial Linear Memory," suggesting a structured, organized way to store visual data rather than just throwing everything into a giant blob of memory. It implies a specific architecture designed for efficient recall based on physical location within the scene.
Meng: From an engineering standpoint, I'm curious how this spatial memory works practically when we're talking about streaming long videos; does it add significant overhead to the processing speed? I need to know if this complexity translates into usable performance gains for our infrastructure.
Lalam: As a language model, I see that the concept of linear memory suggests a very direct path for information retrieval, which is much more efficient than searching through random associations, and this structure could really help in building more coherent long-term representations of an environment.
The paper's summary: Tom: So we've discussed the title, and now we need to get into what LOCI actually does according to the paper. Essentially, this system couples two main memory types: a recurrent state that keeps history compressed and a key-value cache that holds detailed visual information. Jane, can you break down how these two parts work together?
Jane: Absolutely. The summary explains that LOCI uses a hybrid spatial memory architecture where the recurrent memory compresses the entire video history into one fixed-size state, while the key-value cache preserves all the fine visual detail from past observations. This allows it to remember both the general context and the specific appearance of things in a place.
Lu: What I find particularly compelling is how they condition this recurrent memory using projective camera geometry through PRoPE, which means the viewpoint itself gets baked into both what's stored and how we look back at it, giving us precise spatial awareness. This integration of camera geometry with history is something I think could open up really interesting avenues for understanding scene dynamics.
Meng: When the paper talks about the "recurrent context informs these queries under both full and bounded history," that sounds like a crucial mechanism for controlling *when* and *how* the model accesses old information during generation; it’s basically giving the memory a steering wheel to guide its focus.
Lalam: That steering wheel analogy works well; it suggests that instead of just pulling random memories, the recurrent state actively steers the attention towards the most relevant historical data based on where we are looking right now. This structured approach to context management sounds much more intelligent than simple retrieval.
The paper's improvements: Tom: Moving onto what they actually achieved, the summary highlights some specific improvements over existing models. I see they claim LOCI reproduces revisited content more faithfully than other representative world models and even a same-recipe full-softmax model under certain conditions. What are these concrete results?
Jane: The paper points out that LOCI shows an improvement of zero point six two dB on held-out Unreal Engine trajectories when compared to those models, and it achieves a zero point nine nine dB improvement on set A when both models use the same limited set of retained observations for comparison. These numbers show a tangible uplift in how well it reconstructs the scene when revisited after some time has passed.
Lu: I also noticed they mention that the recurrent path has two specific adaptations: PRoPE conditions queries, keys, and values on camera geometry, and then there are token-level delta corrections that refine individual associations. That suggests a very layered approach to error correction, where the system handles both the big picture context and small local details.
Meng: The most practical improvement for me is the claim about streaming long videos at constant memory; they show that under a bounded bank of retained observations, LOCI stays at a constant twenty-three point six GiB memory footprint, whereas full softmax grows to twenty-seven point six GiB when handling longer sequences. That directly addresses our need for efficient long-horizon generation on our hardware.
Lalam: If the model can stream a hundred seconds of video while staying under a fixed memory budget, that capability means we can deploy much more complex, persistent environments without needing massive amounts of dedicated VRAM for every single second of output. This fundamentally alters how we think about deploying large-scale generative models.
Conclusion: Tom: So, to wrap up this discussion on "LOCI: Spatial Linear Memory for Streaming World Models," the main implication seems to be that by combining projectively conditioned recurrent integration with direct historical access, we can create a system that maintains high fidelity when revisiting scenes and allows for efficient streaming. What's your final thought on the paper's overall impact?
Jane: I think the core idea is achieving spatial persistence through this hybrid memory structure, which means models will be much better at maintaining scene integrity over long interactions, and they can do this while being incredibly memory-efficient.
Lu: I believe the ability to condition the recurrent state directly on camera geometry is significant because it ensures that the accumulated context is always spatially relevant to the current viewpoint, which opens up new ways to model dynamic scene navigation and interaction.
Meng: For practical deployment, the constant memory constraint under bounded sparse access is what really moves this paper forward for us; it shows a viable path toward deploying these complex world models in real-world applications without exhausting GPU resources on every single long sequence.
Lalam: I feel that if we can build systems where memory usage scales linearly with the retained context rather than exponentially with the total history length, it sets a new standard for how we design persistent digital environments and how AI interacts with complex visual data.
Ji Xia, Tingting Liao, Xuezhi Liang, Hao Li
Institute of Foundation Models, Mohamed bin Zayed University of Artificial Intelligence
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Comments: 25 pages, 8 figures, 14 tables. Project page: https://xiaji2021.github.io/LOCI/
Project page: https://xiaji2021.github.io/LOCI
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 91/100
The gist: When a camera revisits a previously observed region, a video world model must reproduce what was there before, requiring both remembering past observations and retrieving the right one for the
Key concepts
- Key–Value Cache
- This component acts like a detailed visual library. It stores specific visual information from past observations, preserving fine details of what the camera has seen previously. This allows the model to retrieve exact visual data when revisiting a scene.
- Recurrent Memory
- This is a fixed-size summary that compresses the entire video history into a small state. Instead of storing everything, it keeps a condensed representation of past events, which helps maintain context over long sequences without increasing memory usage linearly.
- Projectively Conditioned Recurrent Memory (PRoPE)
- This mechanism uses the camera's geometric projection to condition both the recurrent memory and attention queries. This ensures that the model understands how different viewpoints relate to each other, allowing the stored history to be relevant regardless of where the camera is currently looking.
Terminology
Summary
When a camera revisits a previously observed region, a video world model must reproduce what was there before, requiring both remembering past observations and retrieving the right one for the current viewpoint. LOCI introduces a hybrid spatial-memory architecture that keeps both representations: key–value caches preserve visual detail while recurrent memory compresses history into a fixed-size state.
The gist
LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model, lowering peak memory at equal length by about 30% relative to full softmax with full history, and streaming long videos at constant memory with a bounded bank of retained observations.
Architecture Overview
LOCI is a hybrid spatial-memory architecture built around the interaction between projectively conditioned recurrent memory and direct historical key–value access. It adopts the layer layout of ARL2, interleaving blocks that combine intra-chunk attention with recurrent memory and blocks that retain historical softmax attention. In each hybrid block, current tokens read the state established by preceding chunks, and a learned gate scales the recurrent readout before it is added to the local-attention output.
Memory Components and Conditioning
The architecture maintains two primary forms of memory:
-
Key–Value Cache: Main attention keeps a key–value cache of past observations, preserving observation-level detail.
-
Recurrent Memory: A fixed-size recurrent state summarizes the entire history at a fixed size (e.g., 15 recurrent + 15 KV blocks). This state is conditioned by projective camera geometry using PRoPE, so
viewpoint enters both memory addressing and stored content.
Hybrid Block Dynamics
The interaction between these memories occurs through a gated mechanism:
** Recurrent context informs these queries under both full and bounded history.
**
The recurrent readout flows into subsequent cache-backed blocks, supplying their queries with accumulated scene context. A learned gate scales the recurrent readout before it is added to the local-attention output, allowing the recurrent state to shape downstream queries.
Historical Access Modes
LOCI supports two modes of historical access:
-
Dense Mode: Historical softmax blocks attend over the evaluated prefix, where main attention KV layers grow with the retained prefix in all historical-attention layers (30 for full softmax and 15 for LOCI).
-
Sparse Mode: This mode retains a recent window and a diverse bank of older views, bounded by a fixed capacity (e.g., 20 bank frames, 8 most recent frames). The bank selection uses a
field-of-view coverage criterion that favors earlier observations adding complementary coverage.
Key Innovations and Results
The paper highlights several key contributions:
** "We develop a hybrid spatial memory that couples projectively conditioned recurrent integration with direct historical KV access, allowing recurrent context to inform historical queries while preserving observation-level detail."**
LOCI improves reference PSNR at revisits by 0.62 dB on held-out Unreal Engine trajectories and by 0.99 dB on set A when both models access the same bounded set of retained observations.
The recurrent path has two adaptations: PRoPE conditions queries, keys, and values on camera geometry, while token-level delta corrections revise individual associations.
With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Performance Comparison
On the public MIND benchmark with a bounded budget, LOCI achieves PSNR +0.89 dB over a same-recipe full-softmax model for 50 segments. Furthermore, on held-out recorded trajectories, LOCI leads by the largest margin in revisit PSNR and exhibits lower LPIPS compared to six of the seven external models evaluated. In the bounded sparse setting, LOCI runs at a constant memory (23.6 GiB) while full softmax grows (27.6 GiB).
Inference Efficiency
LOCI is faster than full softmax in the bounded sparse mode, achieving 5.4 s per video second compared to 8.1 s for full softmax on one H200, as half of its layers attend only within the current chunk. The recurrent state contributes to short-horizon fidelity by shaping queries, and a reset of the recurrent state at every chunk during inference raises local error by +0.021, suggesting its contribution to fidelity.
Conclusion
LOCI combines projective recurrent memory, chunk-level retention, and explicit historical attention in a camera-controlled video model. The architecture keeps both compact spatial context and direct access to retained visual detail, halves the number of main-attention layers that store historical KV, and supports dense and bounded sparse historical access. On the MIND memory benchmark, it has the best MSE, PSNR and SSIM among the world models evaluated. With a bounded bank of retained observations, it streams at constant memory.
Improvements for AI systems
Here are the specific improvements and capabilities enabled by implementing LOCI (Linear Observation-conditioned Interactive) in video world models:
The implementation of LOCI provides significant enhancements across four key areas: fidelity in revisiting scenes, efficiency in long-horizon streaming, robustness under memory constraints, and improved scene understanding during navigation.
Here are the specific improvements and capabilities:
-
Enhancement of Spatial Persistence and Revisit Fidelity
-
Improved Long-Horizon Streaming Capabilities Under Constant Memory Budgets
-
Increased Robustness to Memory Constraints via Bounded Sparse Access
-
Enhanced Scene Understanding through Camera-Conditioned Context Integration
Specific Improvements and Capabilities:
-
The system can reproduce previously observed scenes with high fidelity, even after prolonged absences, by utilizing a hybrid spatial-memory architecture that couples projective recurrent memory with direct historical KV access.
-
The model can stream long videos (e.g., 300 seconds) at a constant memory footprint (e.g., 23.6 GiB under bounded sparse access), significantly lowering the peak GPU memory required compared to full softmax models, which exhaust GPU memory for long sequences (157 s for full history).
-
The system maintains high fidelity during short-horizon revisits (8–20 s) by utilizing a chunk-level retention mechanism that is less sensitive to tokenization noise than per-token retention, leading to lower local error metrics compared to baseline models.
-
The model exhibits superior performance on the MIND memory benchmark, achieving the lowest MSE and highest PSNR/SSIM among evaluated world models under a fixed budget, indicating better scene structure preservation.
-
The camera-conditioned recurrent state is projectively decodable (e.g., maintaining a high R2 score of 0.946 for yaw), meaning the accumulated memory directly informs the current viewpoint and spatial context, leading to more accurate query addressing of historical observations relevant to the current view.
-
The system can selectively access historical observations using a bounded bank of retained views, ensuring that generation continues without growing history storage, while still benefiting from learned scene context derived from the fixed-size recurrent state.
Abstract
When a camera revisits a previously observed region, a video world model should reproduce what was there before. This requires both remembering past observations and retrieving the right one for the current viewpoint. Key-value caches preserve visual detail but grow with video length; recurrent memory is compact but compresses history into a fixed-size state, so individual past observations are no longer directly accessible. We introduce LOCI, a hybrid spatial-memory architecture that keeps both representations. In half of the transformer blocks, main attention keeps a key-value cache of past observations; in the other half, it is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry, so viewpoint enters both memory addressing and stored content. Recurrent readouts flow into subsequent cache-backed blocks and supply their queries with accumulated scene context. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with full history, it lowers peak memory at equal length by about 30% relative to full softmax. With a bounded bank of retained observations, it streams long videos at constant memory and remains more faithful than full softmax under the same budget.
Sources
- AlayaWorld: Interactive Long-Horizon World Modeling - Full Technical Report (v1.1)
- Qwen3-VL Technical Report
- Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models
- Kimi Linear: An Expressive, Efficient Attention Architecture
- Attend Locally, Remember Linearly: Linear Attention as Cross-Frame Memory for Autoregressive Video Diffusion
- Long-Horizon Streaming Video Generation via Hybrid Attention with Decoupled Distillation
- Cameras as Relative Positional Encoding
- Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering
- Compression and Retrieval: Implicit Memory Retrieval for Video World Models
- Advancing Open-source World Models
- Wan: Open and Advanced Large-Scale Video Generative Models
- Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
- Geometry-Aware Implicit Memory for Video World Models
- Echo-Forcing: A Scene Memory Framework for Interactive Long Video Generation
- Addressable Memory for Video World Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models