GeM-NR: Geometry-Aware Multi-View Editing for Nonrigid Scene Changes
cs.CV, cs.AI
Submitted: 2026-06-03
Updated: 2026-09-01
Comments: BMVC 2026. Project page: https://gem-nr.github.io/
Project page: https://gem-nr.github.io
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent developments in multi-view image editing with generative models have brought us a step closer toward general 3D content generation and customization.
Terminology
Abstract
Recent developments in multi-view image editing with generative models have brought us a step closer toward general 3D content generation and customization. Most existing works focus on rigid or appearance-only edits by utilizing the geometry of the unedited scene. This naturally limits these methods to edits that preserve the underlying scene structure. Current nonrigid approaches are limited to object removal and insertion, reflecting the data they are trained on. General nonrigid edits, i.e., edits that substantially and arbitrarily change the scene geometry, remain challenging for existing methods. We propose GeM-NR, a fast and flexible training-free approach for general multi-view consistent image editing, including edits that drastically change the geometry and appearance of the scene. Given an anchor image edited with a chosen 2D editor and a query unedited image, GeM-NR edits the query image consistently with the anchor edit. The method incorporates multiple stages: (i) depth map estimation, where we propose a strategy to maximize the alignment between the 3D point clouds of the edited and unedited scenes, (ii) projection onto a query viewpoint, and (iii) refinement of the obtained image conditioned on the unedited query. We demonstrate the ability of our method to handle edits with significant changes in geometry and appearance, something that existing methods struggle with. We perform an extensive evaluation showing that GeM-NR improves consistency for a wide variety of edit tasks, including generating 3D representations of the edited scene. Both quantitative and qualitative results indicate the state-of-the-art performance of our method in terms of edit quality as well as geometric and photometric consistency across multiple views.
Sources
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- WorldAgents: Can Foundation Image Models be Agents for 3D World Models?
- Guiding Instruction-based Image Editing via Multimodal Large Language Models
- Diffusion-Based Attention Warping for Consistent 3D Scene Editing
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- Efficient-NeRF2NeRF: Streamlining Text-Driven 3D Editing with Multiview Correspondence-Enhanced Diffusion Models
- Edit in 2D, Verify in 3D: Reinforcement Learning for Multi-view Consistent Scene Editing
- Qwen-Image Technical Report
- Towards Scalable and Consistent 3D Editing
- Qwen2 Technical Report
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models