PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion

arXiv:2606.03915 · cs.CV · Submitted 2026-06-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion".

Tom: PatchScene introduces a novel diffusion framework for large-scale LiDAR scene completion that addresses challenges in geometric fidelity, temporal consistency, and computational scalability.

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about "PatchScene: Patch-based Voxel Diffusion for Large-Scale Scene Completion," and the authors are Qingdong Xu, Jiajun Zhu, Shilin Zhu, Xinjing He, Chao Lu, Huanran Wang, Jiyao Zhang. The title tells us right away that they’re using patches and diffusion to tackle big scene completion problems.

Jane: Exactly. The authors are proposing a new way of doing things by using localized patches within a voxel space and applying a diffusion process to generate the geometry in those areas, aiming for consistency across space and time.

Lu: I think the focus on that divide-and-conquer paradigm is smart because it avoids the massive computational cost of processing every single point in one go, which is a problem we see with global representations.

Meng: I wonder how they balance that local generation against maintaining smooth transitions when those patches meet, as that’s usually where artifacts pop up in these kinds of reconstructions.

Lalam: It suggests a path toward building scene understanding models that are more modular, where local details are refined separately before being coherently assembled into the whole picture.

The paper's summary: Tom: So, the paper explains that instead of one giant model looking at everything at once, they partition the entire voxel space into overlapping regular patches and run a diffusion model independently on each patch to clean up and complete just that local area.

Jane: That means they’re essentially treating scene completion like a series of small, manageable tasks—object-level completion per patch—which significantly cuts down the computational load compared to methods that try to process the whole thing simultaneously.

Lu: The forward diffusion process involves adding noise independently to each local patch using an equation like x k t = sqrt t x k zero + sqrt one - t epsilon, and then a neural network predicts the original data based on that noisy input.

Meng: That sounds mathematically sound for generating high-fidelity local geometry, but I need to know how they handle the coordination between those independently generated patches before they get fused back together.

Lalam: The summary shows they introduce two key fusion steps: one to handle spatial overlaps and another to ensure temporal continuity across different frames, which is what really makes this framework unique.

The paper's improvements: Tom: One major improvement they highlight is the spatial fusion mechanism, where they use a stochastic interference method during denoising in overlapping regions to resolve boundary discontinuities that often plague patch-based methods.

Jane: That probabilistic weighting scheme, where a noise prediction has a chance to adopt the global context estimate rather than just its local prediction, is clever for smoothing out those visible seams.

Lu: They also have a temporal fusion step where for the next frame, they use ICP registration to get a transformation and then weight the previous frame's result against the current observation using an adaptively determined scale lambda(p) based on local density consistency.

Meng: That adaptive weighting based on local point density sounds like it could be very useful for dynamic scenes where some areas are dense and others are sparse, ensuring temporal stability where it matters most.

Lalam: And finally, they use an annular-flow diffusion strategy, grouping patches into concentric rings and guiding the generation from the high-density center outwards by conditioning outer ring patches on the inner ones at every timestep.

Conclusion: Tom: So to wrap up, PatchScene introduces a patch-based voxel diffusion paradigm that achieves temporally consistent, high-fidelity completion through local patching, stochastic spatial fusion for boundaries, and an annular flow strategy for infinite spatial extension.

Jane: Basically, they’ve managed to tackle the scale problem by making localized generation tractable and then using smart fusion techniques to stitch those parts together coherently across both space and time.

Lu: The implication is that we can get high-fidelity three dee reconstructions of massive scenes without getting bogged down in the extreme computational requirements of processing everything globally at once <ref:2606.03915#pg0>.

Meng: Practically, this means we could deploy more detailed three dee scene understanding systems in real-world applications where memory and time are strict constraints, rather than just running on smaller, local datasets <ref:2606.03915#pg0>.

Lalam: For culture, I think this framework shows that complexity can be managed through structured decomposition—breaking down a huge problem into localized diffusion tasks with explicit fusion rules is a really powerful pattern for future generative AI.

Qingdong Xu, *Jiajun Zhu*, *Shilin Zhu*, *Xinjing He*, Chao Lu, Huanran Wang, Jiyao Zhang

MEGVII Technology 2Qianli Technology 3Peking University · Northeastern University, China · Northwest Polytechnical University, Xi’an

cs.CV

Submitted: 2026-06-02

Updated: 2026-10-02

Importance score: 90/100

The gist: PatchScene introduces a novel diffusion framework for large-scale LiDAR scene completion that addresses challenges in geometric fidelity, temporal consistency, and computational scalability.

Key concepts

Patch-based Voxel Diffusion
The method discretizes the 3D space into overlapping local patches. A diffusion model is then trained independently on each patch to denoise and complete its local geometry. This avoids processing the entire massive scene at once, making the process computationally feasible by focusing on object-level completion per patch.
Spatial Fusion
This mechanism merges results from neighboring patches to eliminate artifacts at boundaries. It uses a stochastic interference method where predicted noise from individual patches is aggregated into a global noise field, allowing overlapping regions to adopt context from the entire scene context.
Temporal Fusion
To ensure consistency across time, the model combines predictions from consecutive frames. It uses an ICP registration to find movement between frames and then blends the current frame's prediction with the previous frame's result based on a density-consistent weighting factor.
Annular-Flow Diffusion Completion
This strategy handles infinite spatial extent by grouping patches into concentric rings based on distance from the sensor. The diffusion process flows outward from the high-density center to sparse outer regions, allowing distant areas to leverage context from closer, completed parts of the scene.

Terminology

Summary

PatchScene introduces a novel diffusion framework for large-scale LiDAR scene completion that addresses challenges in geometric fidelity, temporal consistency, and computational scalability. The core contribution is a patch-based voxel diffusion paradigm that explicitly generates fine-grained geometry within localized 3D regions while enabling spatially unbounded scene completion through an annular flow strategy.

The gist

PatchScene proposes a novel diffusion-based framework for large-scale LiDAR scene completion that achieves temporally consistent, high-fidelity, and spatially infinite scene completion through a cycle of patch-based completion, fusion, and diffusion.

Patch-based Voxel Diffusion

The method begins by mapping the point cloud space directly into a discrete voxel space to simplify training and improve accuracy. Instead of processing the entire scene at once, the complete voxel space is partitioned into mutually overlapping regular patches with predefined dimensions (height, width, and depth). A diffusion model is then applied independently to each local patch for denoising and completion. This strategy effectively reduces the task to object-level completion per patch, which significantly mitigates the computational and memory burdens associated with traditional voxel-based methods that suffer from cubic growth in computational cost.

The forward diffusion process involves adding noise independently to each local patch:

x k t = sqrt(α¯t x k0 + sqrt(1 − α¯tϵ), t = 1,..., T (2)

The inverse denoising process is handled by a deep neural network trained to predict the original data directly. The model is conditioned on a learnable position encoding pk to ensure the process is sensitive to spatial context, allowing it to recover the estimated noise-free data from the current noisy patch:

x̂ k0 = fθ(x k t, t, pk) (3)

The training objective minimizes the Mean Squared Error (MSE) between the prediction and the ground truth occupancy voxel:

Lpatch(θ) = Ek,t x k0 − ˆx k02 (6)

Patch Spatio-temporal Fusion

To ensure global consistency across spatial and temporal dimensions, PatchScene introduces a mechanism to merge the results from individual patches. This is achieved through two key components:

  1. Spatial Fusion: To resolve boundary discontinuities and noticeable artifacts in the overlapping regions, a stochastic interference mechanism is used during the inverse diffusion process. After obtaining predicted noise for each patch, these predictions are projected and aggregated into a global noise field ˆϵ global. For any point within an overlap zone, the fused noise is generated via a probabilistic weighting scheme:

ˆϵ k fused(p) = B(p) · ˆϵ global(p) + (1 − B(p)) · ˆϵ k(p) (7)

where B is a binary random mask drawn from a Bernoulli distribution, ensuring the noise has a 50% probability of adopting the global context estimate.

  1. Temporal Fusion: This mechanism ensures holistic continuity and geometric consistency across both spatial and temporal dimensions. For frame τ + 1, the model first performs ICP registration to obtain a rigid transformation matrix Tτ→τ+1. The denoised result is then formulated as a weighted combination of the previous frame's prediction and the current frame's observation:

x̂ τ+1 t = λ · x̂ τ t + (1 − λ) · x̂ τ+1 t (8)

The cross-frame guidance scale λ is adaptively determined on a Bird’s-Eye View (BEV) voxel grid based on local density consistency:

λ(p) = min [ρ τ+1(p)/ρ τ(p) + ϵ, 1.0] (9)

Annular-Flow Diffusion Completion

The final stage leverages the physical characteristics of LiDAR scans—being dense near the sensor and sparse far away—to enable infinite spatial extension. The method groups patches into concentric annular regions based on their distance from the sensor's center, denoted as R1 (high-density near range) to RL (distant, sparse region).

The Annular-Flow process proceeds in a center-outward guided generation manner:

For each region Rl, we perform patch-based conditional diffusion sampling to obtain local completions ˆx k0. Critically, the denoising process for patches in Rl is guided by the already completed results from the adjacent inner ring Rl−1 at every timestep t.

This center-outward flow allows high-fidelity information to continuously flow from the core to the periphery, enabling sparse outer regions to leverage semantic context from inner scenes, thus achieving completion for "infinite spatial domains.

Improvements for AI systems

Based on the provided paper, PatchScene offers several key architectural and procedural advancements that can be directly translated into improvements for existing 3D scene completion AI systems:

Here are the specific improvements and what these improved systems can achieve:


The proposed system, PatchScene, improves upon existing point cloud completion methods by integrating a novel diffusion framework with explicit spatial and temporal coherence mechanisms.

  1. Improvement: Explicit Local Voxel Diffusion for Scalability

  2. Mechanism: Instead of processing the entire scene at once (which suffers from cubic computational cost), PatchScene partitions the global voxel space into mutually overlapping local patches. A dedicated diffusion model is then applied independently to each patch for denoising, effectively reducing the problem to object-level completion per patch.

  3. Capability Gained: This allows for high-resolution, fine-grained geometric reconstruction within a tractable computational budget, enabling the system to scale from localized scenes (e.g., 20m range training) to much larger operational ranges (e.g., 50m) without prohibitive memory or time constraints associated with dense global grids.

  4. Improvement: Stochastic Interference Spatial Fusion for Boundary Coherence

  5. Mechanism: To prevent artifacts and boundary discontinuities that arise from fusing independently denoised patches, PatchScene employs a stochastic interference mechanism during the inverse diffusion process within overlapping regions. This mechanism uses a probabilistic weighting scheme (Eq. 7) where the noise prediction of a patch has a 50% chance of adopting the estimate from the global context.

  6. Capability Gained: The resulting system eliminates visible boundary inconsistencies and blurring effects between adjacent patches, ensuring seamless, high-fidelity geometric continuity across spatial boundaries in the reconstructed scene.

  7. Improvement: Confidence-Guided Spatio-Temporal Fusion for Consistency

  8. Mechanism: This mechanism integrates overlapping patches (spatial fusion) with information from preceding frames (temporal fusion). It uses a cross-frame guidance scale, adaptively determined on a Bird’s-Eye View (BEV) voxel grid based on local point density consistency, to weigh the cached prediction from the previous frame against the current frame's observation.

  9. Capability Gained: The system achieves superior temporal consistency and geometric stability across dynamic environments. It can maintain continuous target trajectories and smooth surfaces in sequences by intelligently inheriting structural priors from past frames while adapting quickly to new, sparse information in the current frame.

  10. Improvement: Annular-Flow Diffusion Strategy for Infinite Spatial Extrapolation

  11. Mechanism: Leveraging the physical characteristic of LiDAR data (density decreases radially), PatchScene structures its completion process into concentric annular regions, starting from the dense inner region and progressively guiding generation outwards (Eq. 9). Patches in an outer ring are conditioned on the already completed results of the inner ring at every denoising timestep.

  12. Capability Gained: This enables infinite-space scene completion. The system can reliably extrapolate missing data into distant, sparse regions by propagating high-fidelity information from near-range to far-field regions, a capability crucial for autonomous driving in environments where sensor range is limited but the environment is vast.

In summary, the improved AI system can perform:

  1. High-resolution 3D scene reconstruction that scales efficiently to large outdoor environments.

  2. Seamless generation of geometrically perfect surfaces across spatial boundaries (no artifacts).

  3. Robust, temporally consistent tracking and reconstruction of dynamic objects in sequential sensor data (LiDAR sequences).

Sources

Related papers