3D-Consistent Multi-View Editing by Correspondence Guidance
cs.CV, cs.AI, cs.LG
Submitted: 2025-11-27
Updated: 2026-09-01
Comments: Accepted to LoViF at ECCV 2026
Project page: https://3d-consistent-editing.github.io
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically
Terminology
Abstract
Recent advancements in diffusion and flow models have greatly improved text-based image editing, yet methods that edit images independently often produce geometrically and photometrically inconsistent results across different views of the same scene. Such inconsistencies are particularly problematic for editing of 3D representations such as NeRFs or Gaussian splat models. We propose a training-free guidance framework that enforces multi-view consistency during the image editing process. The key idea is that corresponding points should look similar after editing. To achieve this, we introduce a consistency loss that guides the denoising process toward coherent edits. The framework is flexible and can be combined with widely varying image editing methods, supporting both dense and sparse multi-view editing setups. Experimental results show that our approach significantly improves 3D consistency compared to existing multi-view editing methods. We also show that this increased consistency enables high-quality Gaussian splat editing with sharp details and strong fidelity to user-specified text prompts. Please refer to our project page for video results: https://3d-consistent-editing.github.io/
Sources
- InstantEdit: Text-Guided Few-Step Image Editing with Piecewise Rectified Flow
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Instruct 3D-to-3D: Text Instruction Guided 3D-to-3D conversion
- Diffusion-based Image Translation using Disentangled Style and Content Representation
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space
- One-Step Image Translation with Text-to-Image Models
- Efficient-NeRF2NeRF: Streamlining Text-Driven 3D Editing with Multiview Correspondence-Enhanced Diffusion Models
- Qwen-Image Technical Report
- Towards Scalable and Consistent 3D Editing
- NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models