EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction

arXiv:2507.07410 · cs.CV · Submitted 2025-07-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "EscherNet++: Simultaneous Amodal Completion and Scalable View Synthesis through Masked Fine-Tuning and Enhanced Feed-Forward 3D Reconstruction".

Jane: EscherNet++ proposes a unified diffusion model that simultaneously performs amodal completion and novel view synthesis,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to recap where we are, we’ve seen that EscherNet++ is this diffusion model that tackles both filling in missing parts and creating new views at the same time by using input and feature masking <ref:2507.07410#pg0>.

Jane: That's right, and the paper claims this method is robust because it employs a hierarchical masking strategy involving both input-level and feature-level techniques during training <ref:2507.07410#pg2>.

Lu: The thesis of the paper is that this unified design allows the model to simultaneously perform amodal completion and novel view synthesis, which addresses previous limitations in existing methods <ref:2507.07410#pg1>.

Meng: Essentially, they are proposing a single architecture that handles multiple complex vision problems without needing separate specialized models for each task <ref:2507.07410#pg2>.

Lalam: It seems the core claim is about achieving this unified functionality through this integrated masking strategy during training, which helps the model learn complete geometry from occluded views <ref:2507.07410#pg2>.

Tom: And they specifically point to the creation of a curated dataset from Objaverse-one point zero where silhouettes are overlaid to simulate any possible occlusion, making sure the model learns to predict complete views <ref:2507.07410#pg0>.

Jane: This dataset creation is critical because it directly feeds into the training process, ensuring that when they test it on unseen real-world data, like the OccNVS benchmark, it performs well <ref:2507.07410#pg3>.

Lu: Furthermore, they detail how this model can be queried from any viewpoint to synthesize novel views and seamlessly connect with feed-forward three dee reconstruction models like InstantMesh <ref:2507.07410#pg2>.

Meng: The integration with InstantMesh is a key point because it demonstrates that the model can be used for fast mesh reconstruction, which is a practical application for real-time systems <ref:2507.07410#pg3>.

Lalam: I think the importance here lies in showing that this isn't just about generating pretty images; it’s about creating a system that can reliably handle the complexity of real-world, occluded scenes through unified learning <ref:2507.07410#pg2>.

Tom: Exactly, so the summary is that EscherNet++ is a masked fine-tuned diffusion model designed for zero-shot novel view synthesis with amodal completion ability <ref:2507.07410#pg3>.

Jane: It’s important to remember that they are using both input-level and feature-level masking as the core of their hierarchical strategy <ref:2507.07410#pg2>.

Lu: And their goal is to show that this unified framework can achieve superior performance in either regular or occluded tests <ref:2507.07410#pg2>.

Meng: So, the main point for us is that this paper presents a cohesive system where completion and synthesis happen together efficiently <ref:2507.07410#pg2>.

Lalam: This unified capability suggests that we can build more holistic AI systems capable of understanding scenes in a way that goes beyond just generating isolated images <ref:2507.07410#pg3>.

Conclusion: Tom: So, looking at the authors of this work, we have Xinan Zhang, Muhammad Zubair Irshad, Anthony Yezzi, Yi-Chang Tsai, and Zsolt Kira from the Georgia Institute of Technology and Toyota Research Institute <ref:2507.07410#pg0>.

Jane: It’s interesting how their work bridges the gap between pure generative modeling and practical three dee reconstruction pipelines with models like InstantMesh <ref:2507.07410#pg3>.

Lu: The implication is that the ability to query a model from any viewpoint and integrate it with fast three dee reconstruction without extra training significantly speeds up how we get accurate three dee models <ref:2507.07410#pg2>.

Meng: From an engineering perspective, if this method can genuinely reduce reconstruction time by ninety-five percent while maintaining competitive performance, it makes a big difference in deploying these kinds of tools in live applications <ref:2507.07410#pg3>.

Lalam: I see the broader implication as enabling a future where AI systems can create highly consistent and spatially aware three dee representations of objects from limited or occluded views <ref:2507.07410#pg3>.

Tom: It really feels like EscherNet++ moves us toward a system where complex scene understanding is handled holistically, rather than piece by piece <ref:2507.07410#pg2>.

Jane: The paper demonstrates that the hierarchical masking method, combining input and feature masking, is effective for learning the necessary geometry from occluded inputs <ref:2507.07410#pg2>.

Lu: It shows that by unifying amodal completion and NVS in a diffusion model, we get a much more versatile tool for handling diverse input scenarios <ref:2507.07410#pg3>.

Meng: So, the practical impact seems to be that we can deploy faster three dee reconstruction tools that are inherently more robust to the kind of visual noise and occlusion we see in real-world data <ref:2507.07410#pg3>.

Lalam: For our culture, this advancement means AI can contribute to creating richer digital experiences where geometry and semantics are intrinsically linked, which could influence how we design virtual environments <ref:2507.07410#pg3>.

Georgia Institute of Technology · Toyota Research Institute

cs.CV

Submitted: 2025-07-10

Updated: 2025-07-10

Journal ref: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026, pp. 8846-8856

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 78/100

The gist: EscherNet++ proposes a unified diffusion model that simultaneously performs amodal completion and novel view synthesis, addressing limitations in existing methods by integrating input-level and

Key concepts

Amodal Completion
This is the ability of the model to create a complete, full 3D object from incomplete or occluded input views. Instead of processing one image at a time, EscherNet++ handles multiple masked inputs and synthesizes novel views from different query angles simultaneously.
Novel View Synthesis (NVS)
NVS is the process of generating entirely new, photorealistic images of an object from viewpoints that were not present in the original training data. EscherNet++ performs this while also completing missing parts of the input objects.
Input-Level Masking
This technique trains the model to be robust against occlusions in the initial input views. By randomly masking 50% of an input view during training, the model learns to predict complete views from query poses, significantly improving performance when dealing with missing information in the input.
Feature-Level Masking
This involves randomly masking feature vectors within the encoded feature maps of the diffusion model. The hypothesis is that this forces the model to better understand both semantic meaning and intricate geometric structures, leading to higher quality synthesis and better capture of fine details.

Terminology

Summary

EscherNet++ proposes a unified diffusion model that simultaneously performs amodal completion and novel view synthesis, addressing limitations in existing methods by integrating input-level and feature-level masking for robustness and enabling seamless integration with fast feed-forward 3D reconstruction models.

How it works

The core of EscherNet++ is a masked fine-tuned diffusion model designed to handle occluded input views while synthesizing novel views, moving beyond the limitations of separate stages in existing approaches. The method employs a hierarchical masking strategy, which involves two key aspects:

  1. Curated Dataset: A paired dataset is created from Objaverse-1.0 by overlaying silhouettes onto complete objects with a certain probability to simulate any possible occlusions, ensuring the model learns to predict complete views from query poses.

  2. Input-Level & Feature-Level Masking: The model is fine-tuned using two techniques:

(i) Input-level masking:

The model is trained on a dataset where each input view has a 50 percent chance of being occluded, and the model learns to synthesize novel three other complete views. This technique is the key to robustness to occlusions in input views.

(ii) Feature-level masking:

Inspired by previous works, feature vectors in the encoded feature maps are randomly masked out with a 50 percent chance. The paper hypothesizes that this can further enhance the performance by better understanding of semantics and geometry, as it leads to a better understanding of semantics and better capture of intricate structure details. Empirical results show that the combined hierarchical masking mechanism achieve[s] the best overall performance in either regular or occluded tests.

Simultaneous Completion and Scalability

EscherNet++ is designed to be an end-to-end model, allowing it to simultaneously enables amodal completion and NVS conditioned on a flexible number of input views. This unified design offers advantages such as a smaller storage footprint, enabling faster processing time, and it considers the problem holistically by unifying multi-view amodal completion and NVS tasks. The model's general goal is formulated as a conditional generation problem:

(1) XT ∼ p(XT XR, P R, P T)

where XT and PT represent target novel views and query poses, while TR and PR are input reference views and related poses.

Integration with 3D Reconstruction

The framework is designed to seamlessly integrate with other fast, feed-forward 3D reconstruction models without extra training, thanks to its ability to synthesize views from any given query viewpoint. This capability is demonstrated by integrating models like InstantMesh:

  1. The model synthesizes 36 synthesized views for NeuS-based reconstruction.

  2. An additional 6 views are used for InstantMesh, totaling 42 views for enhanced reconstruction.

This integration allows the performance to be improved to a level on par with SoTA overfitting methods such as NeuS by querying the model from extra viewpoints and providing synthesized views to InstantMesh, achieving a reduction in reconstruction time by 95%.

Performance and Findings

Experiments conducted on occluded zero-shot NVS and 3D reconstruction benchmarks (OccNVS) show that EscherNet++ achieves outstanding performance. Key findings include:

(1) Novel View Synthesis:

The model surpass[es] Zero-1-2-3-based models in terms of overall performance in NVS without considering possible occlusions. Furthermore, it shows superior robustness to occlusions, achieving improvement of at least 5 in PSNR for GSO in all settings over EscherNet when compared to baselines.

(2) Amodal Completion:

Unlike models like InstructPix2Pix and Pix2gestalt which process one image at a time, EscherNet++ takes in varying numbers of masked input images and synthesizes novel views of the object from different query viewpoints including the input viewpoints, standing out for its ability to consider multi-view reference in amodel completion.

(3) 3D Reconstruction:

When integrated with feed-forward models like InstantMesh, the reconstruction performance is enhanced, and the pipeline becomes more robust to occlusions. The method achieves a reduction in reconstruction time by 95% while maintaining competitive performance.

Limitations & Future Directions

The paper identifies two notable limitations:

  1. EscherNet++ exhibits degraded performance on inputs that contain intricate details and complex spatial layouts, where synthesized outputs lack visual and semantic consistency due to the limited resolution (256x256) of both input and output.

  2. It possibly presents failure when presented with out-of-distribution (OOD) occluded inputs, leading to a generation of a flat, unstructured form that resembles a piece of paper instead of plausible completions.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements that can be made to existing AI systems, along with what those improved systems would be capable of:


  1. Improving Novel View Synthesis (NVS) Robustness under Occlusion:

  2. Enhancing 3D Reconstruction Accuracy and Speed via Unified Integration:

  3. Achieving Zero-Shot, Multi-View Amodal Completion in a Single Pipeline:

  4. Developing Scalable and Generalizable 3D Scene Understanding Models:

Here are the specific capabilities of these improved AI systems based on EscherNet++:

  1. Improved NVS Robustness under Occlusion:

  2. Enhanced 3D Reconstruction Accuracy and Speed via Unified Integration:

  3. Achieving Zero-Shot, Multi-View Amodal Completion in a Single Pipeline:

  4. Developing Scalable and Generalizable 3D Scene Understanding Models:

Here is a detailed breakdown of what these improved systems can do:

  1. Improved NVS Robustness under Occlusion:

  2. Enhanced 3D Reconstruction Accuracy and Speed via Unified Integration:

  3. Achieving Zero-Shot, Multi-View Amodal Completion in a Single Pipeline:

Sources

Related papers