AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories".
Jane: The paper was written by Zun Wang, Han Lin, Jaehong Yoon, Jaemin Cho, Yue Zhang et al. from University of North Carolina at Chapel Hill and Nanyang Technological University, Singapore.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Core Problem Summary: Tom: So, let’s look deeper into why those global methods fail and what AnchorWeave identifies as the fundamental issue. It seems like a design flaw that's inherent in the current paradigm.
Jane: The authors explain that even if a small pose or depth estimation error occurs in one view, it can accumulate over time when you try to force all surfaces into one unified global model. This leads to drift, where things just slowly move out of place.
Lu: That means instead of trying to fix one giant geometric mess, they are strategically gathering many clean small pieces of information and stitching them together logically so that the accumulated errors don't matter.
Meng: They describe this as the limitations of global three dee reconstruction; when you’ fuse multiple views, even tiny misalignments cause ghosting or hallucinated content in the rendered anchor videos.
Lalam: Lalam observes that these cross-view artifacts are what destroy visual fidelity, so we're seeing a breakthrough by moving away from a single unified representation and instead focus on local geometric integrity.
The Core Improvement - Retrieval: Tom: That leads directly into the next question: how does AnchorWeave actually improve upon these established methods? It’s not just fixing one giant flaw; it’s fundamentally changing how they access historical context.
Jane: The authors detail that relying on a single global memory is prone to drift because of accumulated errors, as we just discussed. Even if the error is small, the global model tries to force those same surfaces into one unified spot in the three dee space, which causes inconsistencies.
Lu: That implies that instead of trying to fix one giant mess at once, they are strategically gathering many clean small pieces of information and stitching them together logically based on what is actually visible.
Meng: They introduce a mechanism called coverage-driven memory retrieval, which is a huge practical improvement in targeting the necessary data. The system selects specific local memories that haven't been seen yet along the camera path, ensuring we are always gathering new information.
Lalam: And we see the benefit clearly in the results; instead of having those problematic ghosting or drift artifacts from global fusion, we get clean, consistent geometric signals for generation that really support visual clarity.
The Mechanism Deep Dive - Weaving: Tom: This coverage-driven retrieval is a brilliant first step, but the way they utilize all that gathered information is what's truly fascinating—how they manage multiple conflicting inputs.
Jane: But it’s not enough to just gather the memories; you have to use them effectively when generating the frames, which requires a complex control system that can handle multiple sources of geometric guidance.
Lu: It feels like a massive orchestration of information, pulling in specific historical data based on exactly where the camera is looking right now during generation to ensure we are only using relevant context.
Meng: The Multi-anchor Weaving Controller is an elegant solution for fusing these multiple local point clouds into a single, cohesive control signal that the video backbone can understand and use effectively.
Lalam: Lalam finds that this system allows us to be much more precise about what we want to see, ensuring the geometry matches the history perfectly, which is a huge step toward reliable scene generation in any environment.
Conclusion and Outlook: Tom: We’ve seen how AnchorWeave works and exactly what it’s trying to solve; let's wrap up our discussion on World-Consistent Video Generation with Retrieved Local Spatial Memories.
Jane: It’s clear that by moving away from one massive, flawed global memory, we have achieved a major breakthrough in keeping scenes consistent over long periods of time.
Lu: I can only imagine the incredible creative ways this will allow for complex interactions and cinematic storytelling in the future of AI art when things like reliable world-building become standard.
Meng: From an engineering standpoint, it also suggests that scaling up memory management is far more practical than trying to perfect one massive, error-prone three dee reconstruction. It’s a smarter way to build large systems.
Lalam: And I think, by prioritizing local geometric fidelity over global fusion, we are creating a world that is not just visually appealing but structurally trustworthy for the long-term benefit of everyone who experiences it in future AI creations.
University of North Carolina at Chapel Hill · Nanyang Technological University, Singapore
cs.CV, cs.AI
Submitted: 2026-02-16
Updated: 2026-09-04
Comments: Project website: https://zunwang1.github.io/AnchorWeave, Accepted to ECCV 2026
Project page: https://zunwang1.github.io/AnchorWeave
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: AnchorWeave is a memory-augmented video generation framework designed to overcome the persistent challenge of maintaining spatial world consistency in camera-controllable video models.
Key concepts
- Global Three Dee Reconstruction
- This refers to attempting to model an entire scene into one unified 3D space. The hosts note that this approach is prone to drift and artifacts because even tiny misalignments cause inconsistencies when fusing multiple views.
- World-Consistent Video Generation
- The goal of the paper, which aims to create videos where the environment remains structurally trustworthy over long periods. It requires maintaining geometric fidelity and consistency across time, moving beyond simple visual appeal.
- Coverage-Driven Memory Retrieval
- A mechanism introduced by AnchorWeave that selects specific local memories not yet seen along the camera path. This ensures the system is always gathering new, necessary information to maintain accurate context for generation.
- Multi-anchor Weaving Controller
- An elegant control system designed to fuse multiple local point clouds into a single, cohesive signal. It allows the video backbone to use relevant historical data based on the camera's current view, ensuring geometric accuracy.
Terminology
Summary
AnchorWeave is a memory-augmented video generation framework designed to overcome the persistent challenge of maintaining spatial world consistency in camera-controllable video models. Current state-of-the-art methods often rely on fusing historical views into a single global 3D representation, but this approach inevitably introduces cross-view misalignment,
causing artifacts such as ghosting or drift
that degrade the quality of generated content. AnchorWeave addresses this fundamental limitation by replacing the concept of a single misaligned global memory with multiple clean, local geometric memories, allowing it to reconcile their inherent inconsistencies while maintaining strong visual quality and long-term scene consistency.
How Local Memory Replaces Global Consistency
The core insight behind AnchorWeave is that cross-view misalignment artifacts can be mitigated by shifting from a single global 3D memory to per-frame local point cloud memories. Instead of fusing all historical observations into one unified 3D representation, AnchorWeave maintains a collection of independent spatial memory entries, each represented as a local point cloud associated with its original camera pose. This strategy ensures that local point clouds do not accumulate ghosting or drift from multi-view fusion,
providing cleaner geometric signals for rendering-based conditioning.
Coverage-Driven Memory Retrieval
To guide the generation process, AnchorWeave employs a coverage-driven memory retrieval formulation (Section 3.3). Given a target camera trajectory for a specific temporal chunk, the system first constructs a candidate local memory pool by performing an initial Field-of-View (FoV) overlap test. From this pool, it iteratively selects the local memory that maximizes additional visibility coverage
along the target trajectory. This greedy strategy ensures that:
-
The selected memories are compact and complementary, avoiding redundant views.
-
The retrieval process terminates when the visible region is fully covered or the maximum K number of anchors is reached.
Multi-Anchor Weaving Controller
The retrieved local memories are then rendered as multiple anchor videos for conditioning the video diffusion backbone (Section 3.4). These anchors are integrated using a Multi-anchor Weaving Controller, which performs several critical functions:
-
It uses a
shared multi-anchor attention module
to jointly process all K anchor latents, allowing information exchange across anchors. -
It employs a
camera-pose-guided fusion
mechanism that weighs the contribution of each anchor based on its retrieved-to-target relative camera pose, effectively suppressing misaligned or poorly contributing anchors. This ensures that the multiple imperfect renderings are woven into a coherent and spatially consistent output signal injected into the backbone model.
The Iterative Generation Process
AnchorWeave operates in an iterative update–retrieve–generate loop
to achieve long-horizon consistency (Section 3.5). The process is as follows:
-
Update: Starting from a historical frame, local geometry and spatial memory are estimated and initialized.
-
Retrieve: Given the target camera trajectory, relevant local memories are selected using the coverage-driven retrieval strategy (Section 3.3).
-
Generate: The backbone model generates the next video segment under the control of the multi-anchor weaving controller (Section 3.4).
The newly generated frames are then added to the memory bank, enabling AnchorWeave to extend video generation over arbitrarily long trajectories while maintaining strong spatial coherence.
Improvements for AI systems
As a diligent AI researcher, I have thoroughly analyzed the technical architecture of AnchorWeave. The core contribution is not just a new model, but a fundamental shift in how persistent memory is managed and leveraged for long-horizon generation.
The following improvements can be extracted from this research, providing specific blueprints for enhancing existing AI systems (such as Text-to-Video (T2V), Image-to-Video (I2V), and robotic world models).
The Improvement: Replace the reliance on a single, globally fused 3D representation (like a global point cloud) with a bank of per-frame local geometric memories. Each memory entry consists of a local point cloud and its precise camera pose.
What the Improved AI System Can Do:
-
Eliminate Cumulative Drift: The system will no longer suffer from
cross-view misalignment.
Since each local memory is self-contained and not subjected to global fusion errors, the accumulation of small pose or depth estimation errors—which currently causes ghosting and drift in long sequences—is entirely mitigated. -
Maintain High Fidelity over Long Horizons: The system can generate consistent content even when revisiting areas previously seen, maintaining spatial fidelity across multiple segments without the need for a single large-scale, error-prone global reconstruction.
The Improvement: Implement a retrieval mechanism that treats memory selection as a visibility coverage maximization problem, rather than selecting memories based on simple proximity or fixed intervals.
The Improvement: Utilize a Multi-anchor Weaving Controller that uses shared cross-anchor attention and a pose-guided fusion mechanism to synthesize multiple retrieved local memories.
The Improvement: Design the generative process as a continuous, closed loop where newly generated content is immediately converted into local geometric memories and fed back into the memory bank.
By implementing these four interconnected improvements, any AI system transitions from being a short-term visual generator
to becoming a persistent world model.
The resulting systems will exhibit vastly superior spatial consistency, robust handling of complex camera motions (even those outside the training distribution), and the ability to achieve true long-horizon coherence.
Abstract
Maintaining spatial world consistency over long horizons remains a central challenge for camera-controllable video generation. Existing memory-based approaches often condition generation on globally reconstructed 3D scenes by rendering anchor videos from the reconstructed geometry in the history. However, reconstructing a global 3D scene from multiple views inevitably introduces cross-view misalignment, as pose and depth estimation errors cause the same surfaces to be reconstructed at slightly different 3D locations across views. When fused, these inconsistencies accumulate into noisy geometry that contaminates the conditioning signals and degrades generation quality. We introduce AnchorWeave, a memory-augmented video generation framework that replaces a single misaligned global memory with multiple clean local geometric memories and learns to reconcile their cross-view inconsistencies. To this end, AnchorWeave performs coverage-driven local memory retrieval aligned with the target trajectory and integrates the selected local memories through a multi-anchor weaving controller during generation. Extensive experiments demonstrate that AnchorWeave significantly improves long-term scene consistency while maintaining strong visual quality, with ablation and analysis studies further validating the effectiveness of local geometric conditioning, multi-anchor control, and coverage-driven retrieval.
Sources
- Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation
- VRAG: Learning World Models for Interactive Video Generation
- TTT3R: 3D Reconstruction as Test-Time Training
- AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
- VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
- ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
- I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
- Long-Context Autoregressive Video Modeling with Next-Frame Prediction
- Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control
- PreciseCam: Precise Camera Control for Text-to-Image Generation
- InfinityDrive: Breaking Time Limits in Driving World Models
- GS-DiT: Advancing Video Generation with Pseudo 4D Gaussian Fields through Efficient Dense 3D Point Tracking
- Ctrl-World: A Controllable Generative World Model for Robot Manipulation
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- CameraCtrl: Enabling Camera Control for Text-to-Video Generation
- NVComposer: Boosting Generative Novel View Synthesis with Multiple Sparse and Unposed Images
- CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models
- VMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View Memory
- Matrix-game 2.0: An open-source, real-time, and streaming interactive world model
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models
- VisionCoach: Reinforcing Grounded Video Reasoning via Visual-Perception Prompting