Long-Tailed 3D Detection via Multi-Modal Fusion

arXiv:2312.10986 · cs.CV, cs.RO · Submitted 2023-12-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Long-Tailed 3D Detection via Multi-Modal Fusion".

Jane: This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're diving into the specifics of "Long-Tailed three dee Detection via Multi-Modal Fusion" now, and we want to talk about who came up with this work and what that title actually suggests about the research they are doing.

Jane: The title itself tells us immediately that the solution isn't just about detecting three dee things; it emphasizes the multi-modal aspect, meaning they are using more than one type of sensor data together to get better results.

Lu: And then you have that long-tailed detection part, which is the technical challenge they are trying to solve by looking at classes with very different counts across datasets.

Meng: It’s interesting how the authors chose a fusion approach in the title; it suggests a combination method rather than relying on just one sensor or model alone for these tough cases.

Lalam: I think the authors' focus on combining modalities is key because it acknowledges that neither LiDAR nor RGB data alone seems sufficient for reliably catching everything we need to see on the road.

The paper's summary: Tom: We’ve got a good feel for the title, and now let’s look at what the paper actually summarizes, which is where they lay out their main idea about how this problem is being approached in detail.

Jane: They explain that while existing three dee detection methods are improving on common objects, they often miss those rare classes like a stroller or an emergency vehicle because the training data isn't balanced enough.

Lu: The paper points out that the authors are motivated by looking at all annotated classes in datasets like nuScenes, even if it means re-purposing them to see how performance holds up across many different counts per class.

Meng: So, they are essentially saying that for safety, we have to care about those low-frequency events too, which is a shift from only optimizing for the most frequent detections.

Lalam: Their approach centers on proposing a novel Multi-Modal Late Fusion framework that uses independently trained detectors for LiDAR and RGB data to boost performance on both common and rare classes.

The paper's improvements: Tom: Moving into the actual improvements they propose, it seems like they aren't just throwing one big model at the problem; they’ve introduced several distinct algorithmic innovations to tackle this LTthree dee challenge.

Jane: They suggest using hierarchical losses to encourage features to be shared across different classes, which helps ensure that even rare objects get useful features from the training process.

Lu: That idea of training a single feature trunk to predict fine-grained, coarse, and root-level classes is clever because it imposes a structural relationship on how the AI learns about the objects.

Meng: From an engineering viewpoint, that feature sharing sounds efficient because it means we are not wasting computational power by learning entirely separate features for every single rare class.

Lalam: Another key innovation they introduce is Multi-Modal Filtering, which involves fusing detections from a LiDAR-only detector and an RGB-only detector specifically to filter out any detections that are inconsistent between the two sensors.

Conclusion: Tom: We've covered the core ideas and the proposed methods, so now let’s wrap things up by looking at what this whole study means for where we go next in three dee detection research.

Jane: The main conclusion is that their Multi-Modal Late Fusion method significantly outperforms prior work on LTthree dee benchmarks, showing a notable improvement on the six rarest classes from twelve point eight to twenty point zero mAP!

Lu: That result really validates their strategy of matching detections specifically on the 2D image plane rather than trying to match everything in raw three dee space, which they argue is much better for mitigating depth estimation errors.

Meng: It’s good to hear that matching on the projected 2D image plane leads to such a tangible performance gain, especially when compared against end-to-end models that might be trying to fuse everything at once.

Lalam: This research implies that we can build more reliable autonomous systems because they are less reliant on having perfectly balanced training data for every single object type in the real world.

Tom: So, to wrap up on "Long-Tailed three dee Detection via Multi-Modal Fusion," this paper shows that a structured, late fusion approach using modality-specific matching can effectively handle the long tail of rare objects in AV detection. What an exciting development for safety.

Yechi Ma, Neehar Peri, Achal Dave, Wei Hua, Deva Ramanan

Department of Computer Science, Zhejiang University · Robotics Institute, Carnegie Mellon University · Toyota Research Institute

cs.CV, cs.RO

Submitted: 2023-12-18

Updated: 2026-09-30

Comments: Project page: https://github.com/cc50121/lt3d-lf

Code: https://github.com/cc50121/lt3d-lf

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects.

Key concepts

Long-Tailed 3D Detection (LT3D)
This is the challenge of detecting objects in autonomous vehicle datasets where some classes are very common while others, like emergency vehicles, are extremely rare. Existing methods fail to perform well on these infrequent classes.
Hierarchical Losses
These losses help train a single feature trunk that is useful for all object classes simultaneously. Detectors learn to predict three levels of labels: the fine-grained class, the coarse class, and the root-level class.
Multi-Modal Late Fusion (MMLF)
This framework fuses detections from separate LiDAR and RGB detectors after they are independently trained. It specifically matches detections on the 2D image plane to improve accuracy for rare classes by combining calibrated scores probabilistically.

Terminology

Summary

This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects. The authors propose a novel Multi-Modal Late Fusion (MMLF) framework that leverages independently trained uni-modal LiDAR and RGB detectors to achieve state-of-the-art performance across both common and rare classes. This approach is significant because it demonstrates that fusing detections from 2D RGB detectors with 3D LiDAR detections, when matched on the 2D image plane, significantly outperforms end-to-end trained multi-modal detectors.

The Problem: Long-Tailed 3D Detection (LT3D)

The core problem is LT3D, which requires detecting objects from both common and rare classes in real-world scenarios like autonomous driving. Existing benchmarks often ignore rare classes such as emergency vehicle and stroller, despite their importance for safe navigation. The authors motivate the study by re-purposing all annotated classes in datasets like nuScenes to evaluate performance on the entire distribution, which includes few medium many counts per class. To better diagnose performance, they introduce a new diagnostic metric called Hierarchical Mean Average Precision (mAPH), which awards partial credit based on semantic relationships w.r.t to the semantic hierarchy.

Algorithmic Innovations for LT3D

The paper introduces several algorithmic innovations to improve detection results:

  1. Hierarchical Losses: The authors propose hierarchical losses that promote feature sharing across classes by training a single feature trunk, which helps ensure features are useful for all classes. This is implemented by training detectors to predict three labels for each object: its fine-grained class, its coarse class, and the root-level class object.

  2. Multi-Modal Filtering (MMF): To address the issue of sparse LiDAR data struggling with rare objects, they propose an MMF framework that fuses detections from a LiDAR-only detector (for precise 3D localization) and an RGB-only detector (for better recognition). This involves filtering away detections that are inconsistent across modalities.

  3. Multi-Modal Late Fusion (MMLF): The authors investigate three critical design choices for MMLF: whether to train 2D or 3D RGB detectors, whether to match detections in 3D or the projected 2D image plane, and how to fuse matched detections. Their final approach uses 2D RGB detectors, matching on the 2D image plane, and combining score-calibrated predictions with probabilistic fusion.

Key Design Choices in MMLF

The success of the MMLF framework hinges on specific design decisions:

  1. RGB Detector Choice: Experiments reveal that 2D RGB detectors achieve better recognition accuracy for rare classes than 3D RGB detectors. They demonstrate that using 2D RGB detectors, such as DINO, is superior to 3D RGB detectors in the late-fusion context.

  2. Matching Strategy: The authors find that matching on the 2D image plane mitigates depth estimation errors for better matching, and they show this strategy "significantly improves performance for classes with medium and few examples by >10 mAP." They explicitly contrast this with matching detections in 3D space.

  3. Fusion Strategy: The final fusion step involves two crucial components:

List of components:

(1) Score calibration:

The authors calibrate detection confidences per model by tuning a temperature τc for the logit score before applying a sigmoid transform, noting that score calibration and probabilistic fusion notably improves the final performance further.

(2) Probabilistic fusion:

They assume conditional independence given the class label and compute the final score as:

p(cxRGB, xLiDAR) ∝ p(xRGBc)p(xLiDARc)p(c).

Experimental Results and Conclusions

Extensive experiments on the nuScenes benchmark and Argoverse 2.0 datasets reveal that the MMLF approach significantly outperforms prior work for LT3D, particularly improving on the six rarest classes from 12.8 to 20.0 mAP! The results confirm that:

  1. Our simple multi-modal late-fusion method MMLF achieves state-of-the-art LT3D performance on the nuScenes benchmark, significantly outperforming end-toend trained multi-modal 3D detectors.

  2. The framework correctly relabels detections which are geometrically similar (w.r.t size and shape) in LiDAR but are visually distinct in RGB, as illustrated by visualizations of MMLF results.

  3. The approach is robust, achieving an 8.3 mAP improvement averaged over all classes on Argoverse 2 compared to the LiDAR-only baseline (CenterPoint).

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for existing AI systems and what those improved systems could achieve:


  1. Acknowledge and Explicitly Model Long-Tailed Distributions in 3D Detection:

  2. Implement Group-Free Detector Heads with Class-Specific Feature Sharing:

  3. Integrate Semantic Hierarchy Constraints into Training via Hierarchical Losses:

  4. Develop a Multi-Modal Late Fusion (MMLF) Framework for Robust Rare Class Detection:

  5. Employ Modality-Specific Detection Strategies (2D vs. 3D RGB) for Fusion Input:

  6. Utilize Score Calibration and Probabilistic Ensembling During Late Fusion:

Improving AI Systems with These Innovations:

Specifically, the improved AI systems can perform the following tasks:

  1. Detect and accurately localize objects from a comprehensive set of classes, including rare ones (e.g., stroller, emergency vehicle), which are currently poorly recognized by LiDAR-only or end-to-end models.

  2. Improve performance on Few and Medium cardinality classes—the long tail—by leveraging the semantic relationship between objects (e.g., correctly classifying a child as an adult when misclassifying it as a construction worker).

  3. Achieve state-of-the-art Mean Average Precision (mAP) across all classes by utilizing a shared feature backbone that is not constrained by handcrafted grouping heuristics, leading to better generalization for both common and rare objects.

  4. Produce significantly more accurate 3D detections than current end-to-end multi-modal detectors (like BEVFusion or CMT), especially for rare classes, by intelligently fusing high-recall LiDAR data with the superior semantic recognition capabilities of pre-trained 2D RGB detectors (like DINO).

  5. Achieve superior performance when fusing sensor data by matching projected 3D LiDAR detections onto the 2D image plane rather than inflating them to 3D space, which mitigates depth estimation errors and results in a >10 mAP improvement for medium and few-example classes.

  6. Generate final detection predictions with higher confidence by calibrating the scores of RGB and LiDAR detections independently before performing probabilistic fusion, leading to an overall performance boost of 0.9 mAP compared to methods that only fuse matched predictions with semantic agreement.

Abstract

Contemporary autonomous vehicle (AV) benchmarks have significantly advanced multimodal (LiDAR+RGB) 3D detection. However, despite the naturally long-tailed distribution of object classes, existing benchmarking protocols primarily focus on frequent categories (e.g., pedestrian and car), largely overlooking rare but safety-critical classes such as stroller and emergency vehicle. In practice, reliable detection of both common and rare classes is essential for safe autonomous driving. We formalize this problem as Long-Tailed 3D Detection (LT3D), where evaluation encompasses all annotated classes, including rare ones. To address LT3D, we introduce hierarchical losses that promote feature sharing across classes, diagnostic metrics that assign partial credit to semantically reasonable mistakes with respect to the semantic hierarchy (e.g., confusing a child with an adult), and a multimodal late-fusion (MMLF) framework to fuse detections. In particular, we show that rare-class accuracy benefits substantially from MMLF of independently trained uni-modal LiDAR and RGB detectors. Because of the modular design, unlike prevailing end-to-end trained multi-modal detectors that require paired LiDAR-RGB data, MMLF enables the use of advanced unimodal detectors that are trained on large-scale uni-modal datasets with sufficient data for rare classes. Lastly, we examine three fundamental design choices in MMLF, including the RGB detector representation (2D vs. 3D), cross-modal association (3D vs. image plane), and fusion strategy. We find that 2D RGB detectors recognize rare classes more reliably than 3D RGB detectors, image-plane association is more robust to depth estimation errors, and probabilistic score-calibrated fusion consistently yields the best performance. Extensive experiments on nuScenes and Argoverse2 demonstrate substantial improvements of MMLF, establishing a new state of the art.

Sources

Related papers