GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion

summary

Video file (mp4)

The gist

Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios.

In short

GaussianCaR fuses camera and radar data for autonomous driving using Gaussian Splatting as a universal view transformer. It maps both image pixels and radar points into a common Bird's-Eye View (BEV) representation. This method achieves state-of-the-art dense BEV perception while significantly speeding up inference compared to previous models.

Key concepts

Gaussian Splatting (GS)
GS is used as a universal view transformer that efficiently represents 3D data, allowing the model to handle both camera images and radar points. It transforms raw sensor inputs into a dense set of Gaussians that can be easily projected into a unified Bird's-Eye View (BEV) space for fusion.
Pixels-to-Gaussians Encoder
This module processes camera images by first extracting features using an EfficientViT backbone. It then uses a coarse-to-fine strategy to estimate 3D positions for each pixel, determining the size, orientation, and opacity of the resulting Gaussians based on predicted depth distributions.
Points-to-Gaussians Encoder
This module handles radar data using a Point Transformer variant. It processes raw point clouds by creating ordered patches and applying attention mechanisms to model spatial context at the point level. It predicts geometric attributes like position, size, orientation, and opacity for each Gaussian.
Differentiable Gaussian Rasterization
This technique is used to fuse the learned Gaussians from both modalities into a final BEV feature map. It employs an orthographic projection and alpha-blending formula to compute this map, effectively merging the camera and radar information into a single, coherent representation.

Terminology used across episodes

This episode discusses

The paper

GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion · Read on arXiv

Department of Electronics, University of Alcalá, Spain · Department of Cognitive Robotics, Delft University of Technology, The Netherlands

Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios. While vision-only methods have become the de facto standard due to their technical advances, they can benefit from effective and cost-efficient fusion with radar measurements. In this work, we advance fusion methods by repurposing Gaussian Splatting as an efficient universal view transformer that bridges the view disparity gap, mapping both image pixels and radar points into a common Bird's-Eye View (BEV) representation. Our main contribution is GaussianCaR, an end-to-end network for BEV segmentation that, unlike prior BEV fusion methods, leverages Gaussian Splatting to map raw sensor information into latent features for efficient camera-radar fusion. Our architecture combines multi-scale fusion with a transformer decoder to efficiently extract BEV features. Experimental results demonstrate that our approach achieves performance on par with, or even surpassing, the state of the art on BEV segmentation tasks (57.3%, 82.9%, and 50.1% IoU for vehicles, roads, and lane dividers) on the nuScenes dataset, while maintaining a 3.2x faster inference runtime. Code and project page are available online.

DOI: 10.1109/ICRA57385.2026.11697071

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion".

Dev: Robust and accurate perception of dynamic objects and map elements is crucial for autonomous vehicles performing safe navigation in complex traffic scenarios.

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: The title itself, "GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion," tells us exactly what the paper is aiming at: using Gaussian Splatting as a way to fuse camera images and radar points into a Bird’s-Eye View representation. Dev It sounds like they're proposing a new way to handle the view disparity gap between those two different sensors, which is something we always struggle with when trying to make them work together.

Taro: I think the authors are pointing toward using Gaussian Splatting not just for reconstruction, but as a universal transformer mechanism for any sensor modality. Rosa That's what caught my attention; repurposing GS as a view transformer sounds like it could unify how we process entirely different types of input data.

Dev: It seems like the main implication here is that they are creating a framework where both image pixels and radar points are mapped into a shared, sparse three dee space before fusion happens <ref:2602.08784#pg0,both image pixels and radar points>. Taro That sparsity due to the limited number of points in radar clouds is something I'm interested in; how does that constraint affect the final output quality?

Rosa: Well, they propose two specific encoders for this process: a "Pixels-to-Gaussians" encoder for camera data and a "Points-to-Gaussians" encoder specifically for radar point clouds. Dev Having dedicated modules for each modality suggests they are paying close attention to how the input data structure dictates the feature extraction needed.

Taro: That modular approach sounds promising because it allows them to tailor the transformation process precisely to the characteristics of camera images versus raw radar data.

The paper's summary: Rosa: Looking at what they summarize, "GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion" outlines their method as a modality/view to Gaussians to BEV transformation pipeline. Dev So, to put it plainly, the core idea is that they take raw sensor data from both cameras and radar and transform them into a common Bird’s-Eye View representation using three dee Gaussian primitives <ref:2602.08784#pg0,into a common Bird’s-Eye View>.

Taro: That sounds like they are building an intermediary latent space that everyone—both vision and radar data—can understand before the final segmentation happens. Rosa Exactly, they envision this as enabling unified sensor fusion by having dense feature propagation and awareness of uncertainty across different inputs.

Dev: The paper specifically mentions that this approach lets them fuse diverse inputs like pixels and points using these Gaussians, which should lead to more consistent results than traditional fusion techniques. Taro I wonder how they handle the uncertainty aspect you mentioned; does the Gaussian representation inherently carry that information about how reliable a feature is?

Rosa: They do, because they leverage Gaussian Splatting's ability to represent scenes with learnable Gaussians that can be differentiated into planelike representations. Dev That differentiability is key here; it means the entire process, from input to BEV map, can be trained end-to-end efficiently.

Taro: The structure they outline involves two feature encoding branches—the "Pixels-to-Gaussians" and "Points-to-Gaussians"—which are then processed through a Cross-Modal Feature Rectification and Feature Fusion Module. Rosa That fusion stage is where all the magic of combining the camera and radar features into that final BEV map happens.

The paper's improvements: Dev: Now, focusing on their suggested improvements, the paper highlights how they use a multi-scale feature fusion strategy, which includes Cross-Modal Feature Rectification and Feature Fusion Modules. Rosa It seems they are not just relying on one single way to merge the data; they're employing multiple stages to refine that fused representation before decoding it into the final BEV map.

Taro: I see them using a DPT-based decoder after this fusion stage, which suggests a multi-stage transformer architecture is being used to generate the final segmentation maps. Dev That multi-stage approach sounds necessary because you have so much information coming in from two very different sensor types, so you need multiple layers to process it correctly.

Rosa: The main improvement they are proposing is this entire pipeline—the idea of using Gaussian Splatting as a universal view transformer to map raw sensor info into latent features for efficient fusion. Taro That efficiency claim is significant because it suggests a way to achieve dense feature propagation without the computational bottleneck usually associated with fusing high-dimensional data from both sources.

Dev: They also detail how the size, orientation, and opacity of each resulting Gaussian are derived directly from the predicted depth distribution and camera geometry for the camera branch. Rosa And similarly, they use MLP heads in their radar encoder to predict geometric attributes like position and orientation for each predicted Gaussian from the raw point cloud.

Taro: The ablation study points out that adding an early guidance loss and a Dice loss component in the Pixels-to-Gaussians module actually improved performance by one point nine IoU, which shows those specific components are contributing positively to getting better results <ref:2602.08784#pg0>.

Conclusion: Rosa: So, to wrap up what we've heard about "GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion," the paper essentially proposes using Gaussian Splatting as a universal view transformer to map camera and radar data into a common sparse three dee space for efficient fusion <ref:2602.08784#pg0,GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion>. Dev The results show it achieves state-of-the-art performance on dense BEV perception tasks while maintaining an inference runtime that is about three point two times faster than methods like BEVCar.

Taro: I think the biggest implication here is demonstrating a robust and simple framework for effective camera and radar fusion that can operate efficiently in real-world scenarios. Rosa And for me, the real question remains: does this work outside of a highly controlled lab environment? Dev That's the million-dollar question for any system like this; we need to know how stable it is when things go wrong on the road.

Taro: If it can handle misbehaving world scenarios effectively, then having that kind of unified three dee latent representation would be incredibly powerful for autonomous decision-making <ref:2602.08784#pg0>. Rosa It certainly sets a high bar for how efficiently we can integrate these different types of sensory information into a single coherent map for an autonomous vehicle.

Dev: We'll keep an eye on those inference times, Rosa; if it stays at seventy-five point six milliseconds as reported, that makes it much more viable for actual deployment in a car rather than just a research demo.

Taro: Well, the work presented in "GaussianCaR: Gaussian Splatting for Efficient Camera-Radar Fusion" offers a concrete path forward by showing how to use generative techniques like Gaussian Splatting to build efficient, multi-modal perception systems that are better equipped to handle complex traffic environments.

More episodes

← Home