Long-Tailed 3D Detection via Multi-Modal Fusion

summary

Video file (mp4)

The gist

This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects.

In short

The paper tackles Long-Tailed 3D Detection (LT3D) by proposing a Multi-Modal Late Fusion (MMLF) framework for autonomous vehicles. It combines independently trained LiDAR and RGB detectors, matching detections on the 2D image plane, and uses score calibration to achieve state-of-the-art performance across both common and rare objects.

Key concepts

Long-Tailed 3D Detection (LT3D)
This is the challenge of detecting objects in autonomous vehicle datasets where some classes are very common while others, like emergency vehicles, are extremely rare. Existing methods fail to perform well on these infrequent classes.
Hierarchical Losses
These losses help train a single feature trunk that is useful for all object classes simultaneously. Detectors learn to predict three levels of labels: the fine-grained class, the coarse class, and the root-level class.
Multi-Modal Late Fusion (MMLF)
This framework fuses detections from separate LiDAR and RGB detectors after they are independently trained. It specifically matches detections on the 2D image plane to improve accuracy for rare classes by combining calibrated scores probabilistically.

Terminology used across episodes

This episode discusses

The paper

Long-Tailed 3D Detection via Multi-Modal Fusion · Read on arXiv

Yechi Ma, Neehar Peri, Achal Dave, Wei Hua, Deva Ramanan

Department of Computer Science, Zhejiang University · Robotics Institute, Carnegie Mellon University · Toyota Research Institute

Contemporary autonomous vehicle (AV) benchmarks have significantly advanced multimodal (LiDAR+RGB) 3D detection. However, despite the naturally long-tailed distribution of object classes, existing benchmarking protocols primarily focus on frequent categories (e.g., pedestrian and car), largely overlooking rare but safety-critical classes such as stroller and emergency vehicle. In practice, reliable detection of both common and rare classes is essential for safe autonomous driving. We formalize this problem as Long-Tailed 3D Detection (LT3D), where evaluation encompasses all annotated classes, including rare ones. To address LT3D, we introduce hierarchical losses that promote feature sharing across classes, diagnostic metrics that assign partial credit to semantically reasonable mistakes with respect to the semantic hierarchy (e.g., confusing a child with an adult), and a multimodal late-fusion (MMLF) framework to fuse detections. In particular, we show that rare-class accuracy benefits substantially from MMLF of independently trained uni-modal LiDAR and RGB detectors. Because of the modular design, unlike prevailing end-to-end trained multi-modal detectors that require paired LiDAR-RGB data, MMLF enables the use of advanced unimodal detectors that are trained on large-scale uni-modal datasets with sufficient data for rare classes. Lastly, we examine three fundamental design choices in MMLF, including the RGB detector representation (2D vs. 3D), cross-modal association (3D vs. image plane), and fusion strategy. We find that 2D RGB detectors recognize rare classes more reliably than 3D RGB detectors, image-plane association is more robust to depth estimation errors, and probabilistic score-calibrated fusion consistently yields the best performance. Extensive experiments on nuScenes and Argoverse2 demonstrate substantial improvements of MMLF, establishing a new state of the art.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Long-Tailed 3D Detection via Multi-Modal Fusion".

Jane: This paper addresses the challenge of Long-Tailed 3D Detection (LT3D) in autonomous vehicle (AV) benchmarks, where existing datasets focus on common classes and neglect crucial rare objects.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we're diving into the specifics of "Long-Tailed three dee Detection via Multi-Modal Fusion" now, and we want to talk about who came up with this work and what that title actually suggests about the research they are doing.

Jane: The title itself tells us immediately that the solution isn't just about detecting three dee things; it emphasizes the multi-modal aspect, meaning they are using more than one type of sensor data together to get better results.

Lu: And then you have that long-tailed detection part, which is the technical challenge they are trying to solve by looking at classes with very different counts across datasets.

Meng: It’s interesting how the authors chose a fusion approach in the title; it suggests a combination method rather than relying on just one sensor or model alone for these tough cases.

Lalam: I think the authors' focus on combining modalities is key because it acknowledges that neither LiDAR nor RGB data alone seems sufficient for reliably catching everything we need to see on the road.

The paper's summary: Tom: We’ve got a good feel for the title, and now let’s look at what the paper actually summarizes, which is where they lay out their main idea about how this problem is being approached in detail.

Jane: They explain that while existing three dee detection methods are improving on common objects, they often miss those rare classes like a stroller or an emergency vehicle because the training data isn't balanced enough.

Lu: The paper points out that the authors are motivated by looking at all annotated classes in datasets like nuScenes, even if it means re-purposing them to see how performance holds up across many different counts per class.

Meng: So, they are essentially saying that for safety, we have to care about those low-frequency events too, which is a shift from only optimizing for the most frequent detections.

Lalam: Their approach centers on proposing a novel Multi-Modal Late Fusion framework that uses independently trained detectors for LiDAR and RGB data to boost performance on both common and rare classes.

The paper's improvements: Tom: Moving into the actual improvements they propose, it seems like they aren't just throwing one big model at the problem; they’ve introduced several distinct algorithmic innovations to tackle this LTthree dee challenge.

Jane: They suggest using hierarchical losses to encourage features to be shared across different classes, which helps ensure that even rare objects get useful features from the training process.

Lu: That idea of training a single feature trunk to predict fine-grained, coarse, and root-level classes is clever because it imposes a structural relationship on how the AI learns about the objects.

Meng: From an engineering viewpoint, that feature sharing sounds efficient because it means we are not wasting computational power by learning entirely separate features for every single rare class.

Lalam: Another key innovation they introduce is Multi-Modal Filtering, which involves fusing detections from a LiDAR-only detector and an RGB-only detector specifically to filter out any detections that are inconsistent between the two sensors.

Conclusion: Tom: We've covered the core ideas and the proposed methods, so now let’s wrap things up by looking at what this whole study means for where we go next in three dee detection research.

Jane: The main conclusion is that their Multi-Modal Late Fusion method significantly outperforms prior work on LTthree dee benchmarks, showing a notable improvement on the six rarest classes from twelve point eight to twenty point zero mAP!

Lu: That result really validates their strategy of matching detections specifically on the 2D image plane rather than trying to match everything in raw three dee space, which they argue is much better for mitigating depth estimation errors.

Meng: It’s good to hear that matching on the projected 2D image plane leads to such a tangible performance gain, especially when compared against end-to-end models that might be trying to fuse everything at once.

Lalam: This research implies that we can build more reliable autonomous systems because they are less reliant on having perfectly balanced training data for every single object type in the real world.

Tom: So, to wrap up on "Long-Tailed three dee Detection via Multi-Modal Fusion," this paper shows that a structured, late fusion approach using modality-specific matching can effectively handle the long tail of rare objects in AV detection. What an exciting development for safety.

More episodes

← Home