FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "FusionBERT: Multi-View Image--3D Retrieval via Cross-Attention Visual Fusion and Normal-Aware 3D Encoder".
Jane: FusionBERT is a novel multi-view visual fusion framework designed for image–3D multimodal retrieval, addressing limitations in existing methods that focus on single-view alignment and lack structured inter-view feature fusion.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re diving into the paper titled "FusionBERT: Multi-View Image--three dee Retrieval via Cross-Attention Visual Fusion and Normal-Aware three dee Encoder," which basically tells us they've built a framework for linking images from multiple viewpoints to three dee models.
Jane: Exactly, Tom; it’s about moving past methods that only look at one picture and its corresponding three dee shape, suggesting that seeing an object from several angles gives us a much richer understanding of what it actually is.
Lu: The authors are bringing together different ideas, like using a cross-attention mechanism for visual fusion and introducing a specific way to encode the three dee model that considers surface normals.
Meng: Considering the complexity of multi-view input, I'm wondering if their design keeps things manageable when we try to scale it up to handle, say, dozens of views instead of just a few.
Lalam: The focus on integrating visual and geometric information seems like it could significantly improve how our AI systems learn the underlying structure of physical objects across different modalities.
The paper's summary: Tom: Okay, looking at the summary for FusionBERT, it explains that their main goal is to create a unified representation where images and three dee models are strongly aligned in a shared feature space.
Jane: They achieve this by encoding the images into one visual descriptor and the three dee model into a global embedding, then training them so that corresponding features end up very close together.
Lu: The paper details a two-stage attention mechanism for fusing the multi-view images, starting with an L-layer Transformer to let each view contextualize itself before they combine their information.
Meng: That sounds like a lot of computational overhead; I need to know how much training time this two-stage attention process adds compared to simpler pooling methods we might be using now.
Lalam: It seems really important that they are explicitly modeling the interactions between different views, rather than just averaging them together, which is a key distinction in their summary.
The paper's improvements: Tom: The paper highlights two main improvements: first, using cross-attention to capture complementary visual cues across different views instead of just simple pooling.
Jane: That means the system can intelligently figure out which parts of one view are important when looking at another view, which is a lot smarter than just mixing everything up.
Lu: And secondly, they introduce a normal-aware three dee encoder that takes the point cloud as a nine-dimensional vector including position, color, and surface normals.
Meng: Incorporating normals instead of relying only on spatial coordinates and color seems like a very clever way to handle things like textureless CAD models or LiDAR scans where those visual cues are missing or unreliable.
Lalam: By explicitly including surface normal information, the system can capture geometric priors related to local surface orientations and curvatures, which could lead to much more reliable three dee matching in challenging real-world scenarios.
Conclusion: Tom: So, to wrap up on this FusionBERT paper, the main takeaway is that by combining multi-view attention fusion with a normal-aware three dee encoder, they create a way to achieve better image–three dee retrieval than what’s currently available in the SOTA models.
Jane: It really shows how leveraging complementary information from different viewpoints can make the AI representations much stronger and more accurate for matching objects.
Lu: The implications are that we might see a new level of precision in three dee model matching, especially when dealing with complex or ambiguous visual inputs.
Meng: If this approach proves scalable, it could mean we can deploy more robust recognition tools for industrial inspection where texture is often an issue, which is a practical application I'm interested in.
Lalam: For our AI culture, this signals that focusing on the deeper geometric relationships between data points, like normals and view complementarity, leads to genuinely more powerful and versatile models overall.
Wei Li, Yufan Ren, Hanqing Jiang, Jianhui Ding, Zhen Peng, Leman Feng, Yichun Shentu
IROOTECH TECHNOLOGY
cs.CV
Submitted: 2026-04-02
Updated: 2026-09-29
Importance score: 89/100
The gist: FusionBERT is a novel multi-view visual fusion framework designed for image–3D multimodal retrieval, addressing limitations in existing methods that focus on single-view alignment and lack
Key concepts
- FusionBERT
- A novel multi-view visual fusion framework designed for image-to-3D multimodal retrieval. It addresses limitations in existing methods by focusing on structured inter-view feature fusion rather than just single-view alignment.
- Cross-Attention Mechanism
- A two-stage attention mechanism used to fuse multi-view images. It allows each view to contextualize itself before combining information, enabling the system to intelligently capture complementary visual cues across different viewpoints instead of simple averaging.
- Normal-Aware 3D Encoder
- An encoder that takes a point cloud as a nine-dimensional vector, including position, color, and surface normals. This incorporates geometric priors related to local surface orientations and curvatures, which helps in matching challenging inputs like textureless models.
Terminology
Summary
FusionBERT is a novel multi-view visual fusion framework designed for image–3D multimodal retrieval, addressing limitations in existing methods that focus on single-view alignment and lack structured inter-view feature fusion. By introducing a cross-attention-based multi-view visual aggregator and a normal-aware 3D model encoder, FusionBERT aims to create a more robustly fused visual feature for better 3D model matching, establishing a strong baseline for multi-view multimodal retrieval across synthetic and real industrial objects.
The Core Framework
FusionBERT is built upon an extension of the TAMM model (Zhang, Cao, and Wang 2024). The framework aims to learn a unified representation where the multi-view images are encoded into a unified visual descriptor, denoted as fmvimg ∈ R d, and the 3D model M is encoded into a global embedding f3D ∈ R d. The primary goal is to train the model such that these representations are strongly aligned,
meaning semantically corresponding visual and 3D representations are close to each other in the shared feature space.
Multi-View Visual Fusion via Two-Stage Attention
To effectively integrate complementary information from multiple views, FusionBERT employs a two-stage attention mechanism. The first stage uses an L-layer Transformer to contextualize each view with others through self-attention, allowing the aggregator to capture multi-view complementary visual cues while ignoring redundant or ambiguous information among views.
The second stage performs consensus-guided adaptive fusion using cross-attention. This involves computing a multi-view consensus query (q) and then defining projected query (Qh), keys (Kh), and values (Vh) for each attention head, where the cross-attention weights are computed as βh = softmax(Qh(Kh) T / sqrt(dh)). The final fused feature is obtained by concatenating the outputs of all attention heads.
Normal-Aware 3D Model Encoder
The proposed encoder enhances 3D representation robustness by incorporating surface normals into the encoding process. Instead of relying solely on spatial coordinates and color, the input point cloud is represented as a 9-dimensional vector pj = [p posj, p colorj, p normj] ∈ R 9. The point cloud is partitioned into local patches using Farthest Point Sampling (FPS) followed by k-Nearest Neighbor (kNN) grouping. Each patch is embedded through Transformer blocks where the self-attention mechanism models interactions between spatial positions, color appearances, and normals. This explicit inclusion of surface normal information allows the encoder to capture richer geometric priors which are strongly related to local surface orientations and curvatures,
enabling better generalization to textureless CAD models
or LiDAR scans.
Training Pipeline and Performance
FusionBERT utilizes a two-stage training pipeline designed for effective alignment while preserving CLIP’s semantic space. This involves aligning the fused multi-view features with the encoded 3D model features through contrastive learning, which maximizes similarity between matched features and minimizes those for mismatched ones. The model is pre-trained on a large corpus of image–text–3D triplets from datasets like Objaverse (no-LVIS), ShapeNet, ABO, and 3D-FUTURE. Experimental results demonstrate superior performance over State-of-the-Art (SOTA) models like TAMM, OpenShape, ULIP-2, ReCon, and Uni3D under both single-view and multi-view configurations on benchmarks such as Objaverse-LVIS.
Key Contributions
The framework contributes to image–3D alignment by:
-
Modeling
complementary interactions across different views beyond simple pooling
to produce more discriminative representations. -
Improving 3D representation robustness by
incorporating surface normals into the 3D model encoding,
reducing reliance on texture cues, which is a limitation of current encoders like Point-BERT.
The model achieves significant performance gains, showing better image–3D retrieval than SOTA works
and demonstrating the effectiveness of its multi-view visual fusion under variant viewpoints. Furthermore, inference efficiency is competitive, requiring only 0.635 GiB peak memory for a specific configuration of three views and 10,000 input points. The framework also shows consistent improvement in Recall@K as the number of input views increases up to 12. The final implementation successfully achieves Top-1 retrieval
across various tasks, surpassing SOTA models.
Improvements for AI systems
Here are specific improvements that can be made to existing AI systems by leveraging the capabilities of FusionBERT, along with a description of what these improved systems could achieve:
-
A new retrieval system capable of high-accuracy image-to-3D model matching under complex multi-view conditions.
-
A 3D object recognition and retrieval pipeline that is robust to textureless or color-degraded models (e.g., raw LiDAR scans, untextured CAD files).
-
A visual feature fusion module that intelligently aggregates complementary information from multiple object viewpoints, moving beyond simple feature averaging to capture fine-grained inter-view relationships.
Specific improvements and capabilities:
-
The improved system can perform precise 3D model retrieval when a user provides an image captured from three or more distinct viewpoints (e.g., a photograph of a mechanical part taken from several angles).
-
The system can accurately identify and retrieve the correct 3D CAD model or physical object representation even if the input image has poor color information, is grayscale, or lacks surface texture (e.g., retrieving a
textureless
industrial component). -
The visual fusion module will adaptively weigh and integrate features from different camera angles based on their complementary geometric and appearance cues, leading to superior matching accuracy compared to methods relying on fixed pooling strategies.
-
By incorporating surface normals into the 3D encoder, the improved system can distinguish between objects that share similar spatial coordinates but have fundamentally different underlying geometries (e.g., distinguishing a flat plate from a curved one based on local surface orientation).
-
The overall pipeline will achieve state-of-the-art recall rates across diverse benchmarks (like Objaverse-LVIS, ModelNet40, and real-world industrial parts), significantly outperforming existing multimodal retrieval models like TAMM or OpenShape.
Sources
- ShapeNet: An Information-Rich 3D Model Repository
- SparseFlex: High-Resolution and Arbitrary-Topology 3D Shape Modeling
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models