OPERA: Object Perception Enhances Single-view 3D Reconstruction

arXiv:2607.18630 · cs.CV · Submitted 2026-07-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "OPERA: Object Perception Enhances Single-view 3D Reconstruction".

Tom: The gist: Learnt object perception can significantly enhance 3D reconstruction by explicitly injecting perceptual signals from pretrained models into existing reconstruction pipelines to drive more accurate and semantically consistent results.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we’re wrapping up our look at "OPERA: Object Perception Enhances Single-view three dee Reconstruction <ref:2607.18630#pg1,Object Perception Enhances Single-view 3D Reconstruction>." The authors are Huynh and his team. What's the final word on this work?

Jane: They have demonstrated that pre-trained perception models can act as inference-time priors for single-view three dee reconstruction, steering a frozen reconstruction pipeline using this lightweight generation-perception alignment module <ref:2607.18630#pg1>.

Lu: They showed that combining semantic and geometric perception priors yields the strongest results across different pipelines, which is what they found in their experiments <ref:2607.18630#pg1>.

Meng: It’s a neat way to improve fidelity without having to retrain the entire massive reconstruction model from scratch <ref:2607.18630#pg4>.

Tom: So, for the person listening right now, what’s the simplest version of this paper that they should take away?

Jane: The simple thing is that object perception provides a way to inject structured knowledge into the reconstruction process to resolve ambiguity in three dee reconstruction from just one image <ref:2607.18630#pg2>.

Lu: It moves the focus from purely geometric generation toward models that are grounded in understanding what an object actually is <ref:2607.18630#pg4>.

Meng: It suggests a path forward where we can use existing, powerful perception models to give these reconstruction systems a real sense of the object's context <ref:2607.18630#pg2>.

Conclusion: Tom: So, we’re wrapping up our look at "OPERA: Object Perception Enhances Single-view three dee Reconstruction." The authors are Huynh and his team, and they basically show how you can use pre-trained models to guide a reconstruction pipeline without retraining everything from scratch.

Jane: Exactly. It’s about taking those perception models—like vision-language models or depth estimators—and using them as a kind of shortcut to give the three dee reconstruction model better context for what it's looking at.

Lu: What’s really interesting is that they show you can use different kinds of perception, like semantic information from text and geometric information from depth maps, and combining those actually gets the best results.

Meng: From an engineering standpoint, the authors are clever because they don't need to overhaul their existing reconstruction pipelines; it just plugs into them with this alignment module. That’s a practical point for us.

Lalam: I think what this means for culture is that we can start building systems that don't just see shapes, but actually understand the *meaning* behind those shapes, which could change how we design and interact with three dee content in general.

Tom: So, if you boil it down for someone just listening, the main idea is that giving a reconstruction system a strong sense of what an object *is*, based on existing AI perception knowledge, makes the resulting three dee shape much more accurate.

Jane: It’s not about making the math harder; it’s about making the output smarter by injecting that high-level understanding we already have from things like language and depth sensing.

Lu: And they prove that this works across different reconstruction methods, Wonderthree dee and Erathree dee, which shows it’s a general principle rather than just a trick for one specific system.

Meng: The limitation they mention is that the method still has to work with frozen reconstruction models; it’s an enhancement on top of them rather than a replacement for the whole architecture.

Tom: Right, so while it’s not a total overhaul, it’s a really strong way to boost quality quickly by using perception as this powerful inference prior.

Jane: And that leads us perfectly into how those specific perception signals are actually generated and aligned in the technical details we looked at earlier.

Applied Artificial Intelligence Initiative · School of Information Technology

cs.CV

Submitted: 2026-07-21

Updated: 2026-10-08

Project page: https://ynhuhuynh.github.io/perception-3d

Importance score: 76/100

The gist: The gist: Learnt object perception can significantly enhance 3D reconstruction by explicitly injecting perceptual signals from pretrained models into existing reconstruction pipelines to drive more

Key concepts

Object Perception
This refers to the high-level understanding of an object's structure, semantics (what it is), and affordances (how it can be used). In this paper, perception is extracted from models like vision-language models or depth estimation networks. This knowledge acts as a prior, helping the reconstruction model understand what the object should look like in 3D space even with limited input.
Generation-Perception Alignment Module (GPAM)
The GPAM is the core mechanism that bridges perception and reconstruction. It aligns features learned during the image generation process with retrieved perceptual signals. This alignment is achieved by estimating a noise term ($\epsilon$) to steer the generation towards the object's known perceptual understanding, ensuring the output geometry matches both visual appearance and semantic context.
Semantic Perception
This modality uses language descriptions to provide object knowledge. A vision-language model generates a text caption for an input image, which is then encoded into a signal. This signal represents high-level concepts about the object, such as its identity or description, which guides the 3D reconstruction towards semantically correct shapes.
Geometric Perception
This modality leverages depth sensing to provide structural knowledge. Intermediate features from a depth estimation model are used to create a perceptual signal. This signal informs the reconstruction process about the object's actual 3D structure and spatial relationships, leading to better geometric consistency across different views.

Terminology

Summary

The gist: Learnt object perception can significantly enhance 3D reconstruction by explicitly injecting perceptual signals from pretrained models into existing reconstruction pipelines to drive more accurate and semantically consistent results.

Introduction and Motivation

The relationship between object perception and reconstruction is well established in human vision, yet remains underexplored in computer vision The interplay between object perception and object reconstruction is a foundational principle in human vision Humans rely on high-level perceptual understanding, such as object semantics, structure, and affordances, to infer complete 3D shapes from partial observations Modern computer vision approaches to 3D reconstruction largely treat the problem as geometric or generative modelling These methods predominantly rely on data-driven priors or text/image conditioning, and lack an explicit mechanism to incorporate structured perceptual understanding of the object itself Single-view 3D object reconstruction is inherently ill-posed, requiring models to infer unseen geometry and synthesize novel viewpoints from limited input We argue that treating reconstruction purely as a generative or geometric problem is insufficient; instead, it should be grounded in object perception.

Methodology Overview

The proposed framework leverages foundation large-scale models, including vision-language models (VLMs) [30], largelanguage models (LLMs) [31], and foundation monocular depth estimation model [39] to extract perceptual knowledge guiding multi-view image generation The approach is model-agnostic and operates as a plug-and-play enhancement to existing reconstruction pipelines. The core of the method is the generation-perception alignment module (GPAM), which aligns the features learnt from the generation process with retrieved perceptual signals This integration is achieved by incorporating object perception with noise prediction, specifically estimating epsilon (i.e., estimation of ϵ(1:N)) to drive the generation towards the object’s perceptual knowledge.

Perception Modalities and Alignment

The framework demonstrates its effectiveness using two types of perceptual information 1. Semantic perception involves language descriptions, where Qwen2.5VL-7B generates a text caption from the input image y, which is then encoded into a 768-dimensional embedding s to serve as the perceptual signal 2. Geometric perception uses depth sensing via DepthAnything-V2 (ViT-L) [39], where intermediate feature maps are used to produce the depth perceptual signal s, which is observed to lead to better multi-view generation and later 3D reconstruction. The GPAM includes a Feature Aggregation module that aggregates intermediate activations from the U-Net into a compact multi-view representation and a Perception Mapping module that projects these aggregated features to the target perceptual space to produce the aligned feature u.

Training and Loss Functions

The training process involves minimizing a loss function L(s, u) which measures the similarity between perceptual signals s and generation-perception aligned features u For the semantic modality, the loss is defined as L(s, u) = 1 − coss(u) For the geometric modality, a loss based on similarity between generation-perception aligned features and depth features is used This geometric loss is decomposed into terms involving L2-std and Centered Kernel Alignment (CKA) losses for the front, front-left, and front-right views. The CKA loss is applied to the side views to encourage cross-view geometric consistency without requiring exact depth alignment as in the front view The training is performed by freezing both the multi-view image generative model (the U-Net) and the perception model(s) while training the GPAM for each perception modality.

Experimental Results

Experiments were conducted with two state-of-the-art single-view 3D reconstruction pipelines: Wonder3D [19] and Era3D [14] The results show that both semantic and depth knowledge individually reduce Chamfer Distance (CD) over the baseline, with their combination yielding the strongest result, such as a −9.0% reduction in CD for Wonder3D + semantic + depth The usefulness of perception for reconstruction is transferred across architectures, where combined perception guidance yields the lowest CD among all configurations Qualitative results show that the method recovers fine geometric details and corrects shape collapse on challenging objects.

Conclusion

We have shown that pre-trained perception models can serve as inference-time priors for single-view 3D reconstruction, steering a frozen single-view 3D reconstruction pipeline through a lightweight generation-perception alignment module Across two architecturally distinct pipelines, Wonder3D and Era3D, semantic and geometric perception priors each improve reconstruction quality, and combining them yields the strongest result The framework is model-agnostic and, in principle, extensible to other generative settings where pre-trained perception models exist and intermediate features admit gradient-based alignment. This work motivates further research on perception as a general prior for generative systems.

A Detailed Architecture Diagrams

Figure 3 illustrates the architecture of our method applied to Wonder3D with the semantic modality, showing how the Feature Aggregation module aggregates U-Net activations and the semantic Perception Mapping module produces the aligned feature u compared against caption embedding ss Figure 4 details the architecture applied to Wonder3D with the geometric (depth) modality, where aggregated RGB- and normal-stream features are mapped to u compared against DAv2 F3 target sd

A.1 Era3D Architectural Adaptations

The framework applied to Era3D shares the same conceptual structure as with Wonder3D, except for lighter aggregation width and unified forward pass The aggregation channel width is backbonedependent: Cagg = 384 for Wonder3D and Cagg = 128 for Era3D, whose Feature Aggregation module is deliberately lighter In both backbones, the geometric Perception Mapping module receives the channel-wise concatenation of the colour- and normal-domain aggregated features for the three supervised views (front, front-left, and front-right) The two backbones differ in how the two domains are produced: Wonder3D generates colour and normal views through two separate U-Net streams, whereas Era3D produces both domains in a single unified U-Net forward pass over a batch-concatenated 2N-view input

B Pseudo-code

Algorithm 1 presents the training code of our method, where the GPAM parameters γ are updated by minimizing L(s, u) Algorithm 2 presents the inference code, where noise prediction is guided by gradients derived from L(s, u)

C Semantic Perception Ablation

The two-step VLM+LLM approach consistently outperforms the single-step VLM-only baseline across all tested guidance weights Text captions provide a cleaner semantic signal because generating an intermediate caption encourages the VLM to summarise the object into higher-level concepts before re-encoding with T5

D Depth Perceptual Target

We select F3 for the depth perceptual signal s because it requires only a mild 1.

Improvements for AI systems

  1. Begin single-view 3D object reconstruction by explicitly injecting perceptual signals from pretrained perception models capturing semantic and geometric information to drive the reconstruction of an object from its single image, which resolves ambiguities that classical and learning-based methods focus on mapping images to shapes using implicit representations or learnt priors often struggle with.

  2. Enable model-agnostic integration by providing a plug-and-play manner enhancement, allowing the framework to be applied to existing pipelines such as both feed-forward reconstruction networks and generative approaches.

  3. Enhance semantic consistency by leveraging language descriptions, specifically employing a two-step combination of a VLM (i.e., Qwen2.5VL-7B) and an LLM (i.e., T5) to extract a guidance signal that is less tied to low-level appearance than single-step VLM features, leading to better standalone reconstruction performance.

  4. Improve geometric fidelity by utilizing depth sensing as a modality, where the loss function is structured as LL2std(front) + ωfront-right LCKA(front-right) + ωfront-left LCKA(front-left) with an optimal weight of w = 0.50, ensuring the model learns structural similarity in a viewpoint-invariant manner.

  5. Achieve robust, combined reconstruction by leveraging multiple perceptual signals where combining semantic and geometric perception yields the strongest result, demonstrating that these modalities supply non-redundant corrections to the generative prior.

Sources

Related papers