FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

arXiv:2605.29997 · cs.CV · Submitted 2026-05-28 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views".

Jane: FRUC presents a feed-forward 3D Gaussian Splatting framework designed for dynamic scene reconstruction from uncalibrated collaborative driving views,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, FRUC tackles the problem of dynamic scene reconstruction from uncalibrated collaborative driving views by using a feed-forward three dee Gaussian Splatting framework that avoids needing precise spatial calibration <ref:2605.29997#pg0,dynamic scene reconstruction from uncalibrated collaborative driving views>. Jane, what do you think about the title itself?

Jane: I think the title really nails what they are doing: reconstructing a dynamic scene using collaborative, uncalibrated inputs through a single pass. It’s very descriptive of the core problem they are solving.

Lu: The authors tackle a fundamental bottleneck in current multi-agent reconstruction frameworks by aiming for one-shot inference without relying on precise multi-agent calibration, which is where the creativity really shines.

Meng: I'm interested in how they handle that calibration issue because, practically speaking, needing perfect alignment every time you start a new scene is just not feasible in a car setting.

Lalam: It speaks to an AI culture where systems can operate fluidly and adapt to real-world scenarios without being tied down by rigid pre-set configurations.

The paper's summary: Tom: Moving into the details, the paper summarizes FRUC as a framework that uses a visual grounded geometric Transformer backbone to handle these flexible multi-vehicle views, which then process those inputs using a DINO encoder and subsequent augmentations to create spatio-temporal anchors.

Jane: That sounds complex, but simply put, they take all the different images from the vehicles, turn them into tokens that have identity and time information attached to them, and feed those into a transformer that understands the geometry of what's happening across different views simultaneously.

Lu: The core mechanism they propose is using an ego-centric causal occlusion field to explicitly track how occlusions evolve over time, which is a sophisticated way to model the dynamic nature of blind spots.

Meng: That tracking ability sounds powerful, especially when it’s coupled with a cross-agent latent residual denoising process that ensures the ego vehicle's observations stay reliable while still using the other vehicles' data.

Lalam: I find that ability to selectively use collaborative information based on occlusion is really interesting; it suggests a more intelligent way for AI systems to integrate external data rather than just dumping it all in.

The paper's improvements: Tom: Now we're talking about what makes FRUC different, and the authors highlight two main improvements: first, introducing that ego-centric causal occlusion field to capture occlusion evolution, and second, the cross-agent latent residual denoising process.

Jane: The improvement in terms of the occlusion field is that it doesn't just see a static blind spot; it models how those occlusions change as objects move around, which is crucial for dynamic driving scenarios.

Lu: The cross-agent latent residual denoising module is clever because instead of just mixing features together blindly, it uses the occlusion prior to explicitly modulate feature supplementation, preventing uncalibrated geometric drift from messing up the ego vehicle's visible parts.

Meng: That modulation sounds like a necessary safeguard; if you let uncalibrated data corrupt your primary view, the whole reconstruction becomes useless for safety-critical tasks.

Lalam: This points toward a future where collaborative AI can handle uncertainty much better, allowing it to trust its own sensors when collaboration introduces noise or drift.

Conclusion: Tom: So, wrapping up this discussion on FRUC: the paper shows that by combining a visual grounded geometric Transformer with these specific prior modules, they establish a new state-of-the-art in both reconstruction quality and efficiency for collaborative driving views.

Jane: Exactly, the implication is that we can achieve very high fidelity three dee models of dynamic scenes from multiple vehicle inputs without needing tedious per-scene optimization or perfect initial calibration <ref:2605.29997#pg0>.

Lu: The future work they suggest seems to point toward extending this capability into more complex what-if analysis and perhaps integrating these scene reconstructions directly into continuous closed-loop simulators for training generalizable world models.

Meng: From a practical standpoint, achieving inference speed of zero point seven seven seconds while maintaining this level of accuracy is the key; it means this technology could actually be used in production systems right away.

Lalam: This advancement means that AI systems will be much more capable of understanding complex, shared physical spaces, which really elevates the potential for intelligent automation everywhere we look.

Yihang Tao, Yu Guo, Zhengru Fang, Haonan An, Yuguang Fang

City University of Hong Kong

cs.CV

Submitted: 2026-05-28

Updated: 2026-10-02

Importance score: 83/100

The gist: FRUC presents a feed-forward 3D Gaussian Splatting framework designed for dynamic scene reconstruction from uncalibrated collaborative driving views, overcoming the limitations of existing methods

Key concepts

Visual Grounded Geometric Transformer (VGGT)
This is the core backbone that enables one-shot reconstruction from flexible multi-vehicle views. It uses a DINO encoder to process images and then employs an Alternating-Attention mechanism to aggregate contextual information across different frames and agents, inferring the 3D geometry of the scene.
Ego-Centric Causal Occlusion Field (COF)
This module models how blind spots change over time. It first estimates a base spatial footprint using a dynamic head and then uses an Agent-wise Causal Masked Attention layer to track motion, predicting how the object's occlusion area adaptively expands along its path.
Cross-Agent Latent Residual Denoising (CALRD)
This process aligns features from different agents by using the occlusion prior to guide feature supplementation. It predicts a correction term that is added to mixed tokens, ensuring that ego-only observations are protected from uncalibrated drift when collaborative features are incorporated.
Prior Stabilization (Lreg)
This auxiliary loss anchors the occlusion mask to a fixed base footprint. Its purpose is to enforce the principle that the ego vehicle should rely on its own high-fidelity observations for static, visible areas, preventing collaborative data from corrupting local geometry.

Terminology

Summary

FRUC presents a feed-forward 3D Gaussian Splatting framework designed for dynamic scene reconstruction from uncalibrated collaborative driving views, overcoming the limitations of existing methods that require precise spatial calibration and slow per-scene optimization. The core contribution is a novel approach that conceptualizes distributed multi-vehicle networks as an unstructured egocentric system, enabling seamless blind-spot completion by leveraging ego-centric occlusion priors and deterministic residual denoising to preserve reliable ego observations while robustly assimilating uncalibrated cross-agent information.

How it works

The framework is built upon a visual grounded geometric Transformer (VGGT) backbone to enable one-shot, calibration-free inference from flexible multi-vehicle views. To process the input sequence of uncalibrated multi-agent images, the system first flattens and tokenizes them using a DINO encoder. These tokens are then augmented with learnable agent identity embeddings and temporal embeddings to create spatio-temporal anchors, resulting in augmented tokens that serve as inputs to the VGGT Alternating-Attention (AA) feature backbone. This backbone aggregates inter-frame and intra-frame contexts to infer implicit multi-view geometry, producing aggregated deep features where uncalibrated collaborative perspectives interfere the ego’s geometric representation.

Ego-Centric Causal Occlusion Field

To model the continuous evolution of blind spots caused by dynamic objects, FRUC introduces an ego-centric Causal Occlusion Field (COF). This field explicitly captures occlusion evolution by first estimating a base dynamic probability map using a DPT-style dynamic head to identify the current spatial footprint. To capture kinematics, it employs an Agent-wise Causal Masked Attention layer to extract a latent kinematic field, formalized as motion evolution tokens. Based on this field and the base footprint, an Occlusion Decoder models the occlusion evolution by predicting a spatio-temporal expansion residual that dictates how the object’s footprint adaptively expands along its motion trajectory to cover the dynamically occluded background. This yields a structured spatial prior, denoted as the Occlusion Uncertainty Prior.

Cross-Agent Latent Residual Denoising

Cross-agent feature alignment is formulated as a constrained latent residual denoising process via the Cross-Agent Latent Residual Denoising (CALRD) module. Instead of naive feature concatenation, CALRD utilizes the occlusion prior to explicitly modulate feature supplementation, ensuring that collaborative features may introduce uncalibrated geometric drift into visible regions. The process involves extracting clean ego-only features as a structural reference and feeding them, along with the mixed tokens and occlusion prior, into a dual-branch convolutional architecture. This architecture predicts a residual correction term that is added to the noisy mixed tokens: Fˆ τ m,k = F˜ τ m,k + Zout(Φf ea(F˜ τ m,k) + Cτ pri,k). This mechanism ensures that for unoccluded regions (where Mτ occ ≈ 0), the zeroconvolutions suppress collaborative updates to protect the ego vehicle’s observations.

Model Training and Auxiliary Losses

Training is conducted via a two-stage progressive curriculum. Stage I involves single-agent pre-training to establish robust ego-centric reconstruction, training the Gaussian rendering head and dynamic head using only single-agent sequences. Stage II introduces cross-agent adaptation by freezing the prior modules and training the FRUC modules with multi-agent sequences, guided by auxiliary objectives:

  1. Prior Stabilization (Lreg): Anchors the occlusion mask to a dilated base footprint to instill a conservative principle, ensuring the ego vehicle should fundamentally trust its local high-fidelity observations for static, visible regions.

  2. Cooperative Completion (Lcoo): Enforces photometric penalties exclusively within the dynamic occlusion boundary to force the network to expand Mτ occ,k to fetch and align features from collaborative agents for reconstruction.

  3. Ego-Manifold Regularization (Lden): Applies a global Mean Squared Error objective on the latent space between the denoised features and ego-only features, acting as a structural prior to prevent catastrophic forgetting of ego capabilities during multi-agent fine-tuning.

Evaluation and Results

FRUC is evaluated on V2XReal and UrbanIng-V2X datasets using metrics including PSNR, SSIM, LPIPS for full image reconstruction, Dynamic-only settings, and NIQE for blind-spot completion. Extensive evaluations show that FRUC establishes a new state-of-the-art by significantly outperforming existing methods in both rendering quality and efficiency. Specifically, the framework achieves superior performance in Blind-Spot Completion, demonstrating superior blind-spot complemention capability compared to baselines, while maintaining a high inference speed of 0.77s, which is "significantly faster than optimization-based paradigms.

Improvements for AI systems

Based on the provided research paper, here are specific improvements for AI systems derived from FRUC (Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views), along with what those improved systems can achieve:


  1. Improved 3D Scene Understanding and Geometry Completion in Uncalibrated Multi-Agent Scenarios.

  2. Enhanced Real-Time What-If Analysis for Autonomous Driving Simulation and Decision Making.

  3. Robust, Calibration-Free Collaborative Scene Editing and Novel View Synthesis in Dynamic Environments (Blind Spot Filling).

Specific capabilities of these improved systems:

  1. The system can accurately reconstruct the 4D (spatio-temporal) geometry of a driving scene by fusing uncalibrated views from multiple vehicles without requiring precise, hardware-level spatial calibration or auxiliary sensor data (like LiDAR point clouds) for initialization.

  2. It can seamlessly and reliably complete blind spots in the ego vehicle's perspective—the areas physically occluded by other dynamic agents—by leveraging the complementary information from collaborating vehicles, resulting in a geometrically complete 3D environment.

  3. The system can generate high-fidelity novel views (novel view synthesis) of any arbitrary viewpoint within the reconstructed scene, including synthesizing what-if scenarios by editing objects or removing elements to test safety-critical decisions, all while maintaining high rendering quality and inference speed (achieving state-of-the-art performance on V2XReal and UrbanIng datasets).

Sources

Related papers