JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space".
Jane: JointEdit3D introduces a feed-forward framework for 3D scene editing that operates within a unified RGB-geometry reconstruction-generation latent space,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Moving on to the title and authors of "JointEditthree dee," it’s all about fusing feed-forward techniques with a unified latent space for three dee scene editing. The team behind this work is Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong, Jiachen Xu, Yuan Xie, and Xin Tan. This paper shows how they're adapting an existing RGB-geometry reconstruction-generation latent space specifically for editing scenes.
Jane: It’s interesting how they’re building on a latent space that already handles both the visual and the three dee parts together; it suggests that the structure of the scene is already encoded in a way that makes editing geometry easier than in decoupled systems.
Lu: The authors are clever because they are not starting from scratch; they’re leveraging established concepts like voxel-structured three dee latents which have shown promise for object-level control, but they're applying it to the more complex task of scene-level editing.
Meng: I wonder if that existing latent space is flexible enough to handle the kinds of complex edits we see in real applications without needing a complete overhaul of the underlying architecture.
Lalam: If this unified latent space can effectively merge appearance and geometry, it implies a more intuitive way for an AI to understand scene editing because it doesn't have to keep two different representations in mind separately.
The paper's summary: Tom: To summarize the core of "JointEditthree dee," they propose a feed-forward framework that performs asymmetric latent inpainting. Essentially, they observe only one edited RGB reference frame and then the model generates everything else—the rest of the RGB views and the geometry—conditioned on that single reference and any language instructions you give it.
Jane: That "asymmetric latent inpainting" is key here; it means they are not trying to reconstruct an entire scene from scratch based on a prompt, but rather they are using what we already have—the edited frame—as the known observation to guide the generation of the rest of the scene and its structure.
Lu: The core idea is that edit decisions, appearance synthesis, and geometry updates should all happen within that single latent state simultaneously, which is a departure from previous methods where those steps were separated in time or space.
Meng: So they’re treating editing as an inpainting problem within this shared space; it simplifies the pipeline by making appearance synthesis and geometry prediction a single step rather than two sequential ones.
Lalam: This approach to latent inpainting, conditioned on both the reference frame and language context, suggests that future AI could perform scene manipulation with much more contextual awareness, understanding *why* an edit was made based on what we are told.
The paper's improvements: Tom: The authors introduce a few specific technical improvements to make this work better. First, they have this dedicated SceneAnchor Branch that injects source-scene features back into the main generator as residual corrections to keep the original content intact without forcing a direct copy.
Jane: That SceneAnchor Branch sounds like a clever way to handle preservation; instead of using an explicit mask, it acts like an implicit localizer, adapting the source information dynamically based on where the edit is happening.
Lu: I think that mechanism is important because it addresses the difficulty of preserving unedited content when you're optimizing for a specific edit region within a shared latent space. It’s about making sure the source scene structure stays grounded while allowing targeted changes to happen.
Meng: From an engineering standpoint, injecting features as residuals instead of relying on a static mask simplifies the training process and makes the model more adaptive to different editing scenarios during inference.
Lalam: If we can achieve that level of implicit localization without needing manual masks, it could significantly improve how we build cultural models because we wouldn't need to pre-define every region that must remain unchanged.
Conclusion: Tom: So, to wrap up "JointEditthree dee," the implication is a feed-forward approach that couples appearance synthesis and geometry prediction within a unified latent space for scene editing, leading to results on SceneEditthree dee-15K and SceneEditthree dee-Bench that show improved edited-region quality while maintaining decent background metrics.
Jane: Exactly; they show that by using edit/background-aware losses, they can separate the changes from the stable parts of the scene in both RGB and geometry, which helps keep things looking right overall.
Lu: This work pushes the idea that we don't need sequential optimization or cascaded pipelines anymore when aiming for scene-level editing; we can just generate what’s edited as a single target.
Meng: It’s impressive that they managed to achieve these quality gains while still maintaining inference times that are substantially faster than feed-forward baselines under the full-three dee protocol.
Lalam: Ultimately, this framework suggests a path toward AI systems that can perform complex three dee manipulation with better structural fidelity and scene coherence guided by a single reference point.
Xinnan Zhu, Ruijie Xu, Jiayu Ying, Daoguo Dong, Jiachen Xu, Yuan Xie, Xin Tan
East China Normal University · Shanghai Artificial Intelligence Laboratory (Shanghai) · Fudan University
cs.CV
Submitted: 2026-06-11
Updated: 2026-09-29
Project page: https://xinnan-zhu.github.io/JointEdit3D-Page
Importance score: 88/100
The gist: JointEdit3D introduces a feed-forward framework for 3D scene editing that operates within a unified RGB-geometry reconstruction-generation latent space, addressing the limitations of existing methods
Key concepts
- Unified Latent Space
- This is a shared space that already encodes both visual (RGB) and three-dimensional geometry information together. The paper leverages this existing structure, suggesting that scene structure is already encoded in a way that makes editing geometry easier than in systems with separate representations.
- Asymmetric Latent Inpainting
- This technique involves observing only one edited RGB reference frame and using it to generate the rest of the scene, including other RGB views and geometry. It uses the known edited frame as guidance to reconstruct the unedited parts of the scene based on language instructions.
- SceneAnchor Branch
- This is a technical improvement that injects features from source scenes back into the main generator as residual corrections. This helps keep original content intact during editing by acting like an implicit localizer, adapting source information dynamically without needing explicit masks.
Terminology
Summary
JointEdit3D introduces a feed-forward framework for 3D scene editing that operates within a unified RGB-geometry reconstruction-generation latent space, addressing the limitations of existing methods by coupling appearance synthesis and geometry prediction during editing. This approach moves beyond decoupled optimization or cascaded pipelines, aiming to perform scene-level editing by treating edited appearance and geometry as a shared generation target. The framework is supported by a novel dataset, SceneEdit3D-15K, and benchmark, SceneEdit3D-Bench, providing the necessary paired resources for standardized evaluation of 3D scene editing fidelity.
Unified Latent Space Formulation
The core idea is to adapt a unified RGB-geometry reconstruction-generation latent space (formed by Gen3R) to 3D scene editing. This space jointly models appearance and geometry, where the full latent is defined as a concatenation of RGB and geometry components:
z(V) = [z rgb; z geo] ∈ R C×F ×H×2W, where the first and second spatial halves store RGB and geometry under matched indices.
The model predicts the complete edited RGB-geometry latent by treating editing as RGB-geometry latent inpainting,
where a single edited reference frame serves as the known observation, and the model generates all remaining views and geometry positions.
Asymmetric Latent Inpainting
JointEdit3D performs asymmetric latent inpainting, meaning it observes only a single edited RGB reference latent while generating the rest of the scene. This is achieved by:
-
Encoding the edited reference frame into a specific temporal position within a zero-filled tensor, creating an observed latent frame and an associated binary mask.
-
Concatenating this observed latent with the noisy target latent before passing it to the transformer, conditioning it on both the edited reference and language instructions (context).
Source-Scene Preservation via SceneAnchor Branch
To ensure that edits preserve source content outside the specified region without forcing direct copying, JointEdit3D introduces a dedicated SceneAnchor Branch. This branch receives edit cues from the main branch and injects source-scene features back as residual corrections:
This branch acts as an implicit edit-region localizer without requiring a mask.
The anchor block is interleaved with selected layers of the main generator, allowing source conditioning to adapt to the current edit state rather than acting as a static copy signal.
Edit-Aware Latent Losses
To balance edited-region fidelity with unedited-content preservation, JointEdit3D employs edit/background-aware losses instead of a standard Mean Squared Error (MSE) objective. These losses are designed to separate edited and background regions in both RGB and geometry latents:
-
The loss is decomposed into edit and background regions for each modality (RGB and geometry).
-
Positions are upweighted based on two factors: positions farther from the reference frame, where the latent change is larger, or positions within the intended edit region.
-
A linear weight function, such as w h i = 1 + αd h i, is applied inside the edit region to give larger gradients to regions with larger latent changes.
Evaluation and Datasets
To address the lack of paired resources, JointEdit3D introduces SceneEdit3D-15K (a dataset with 15K paired editing samples and renderer-provided 3D annotations) and SceneEdit3D-Bench (a curated 100-sample benchmark). These resources allow for standardized evaluation across various edit types, including Add,
Delete,
Move,
and Appearance Change.
Experiments show that JointEdit3D improves edited-region quality and 3D structural completeness over prior baselines while maintaining competitive background preservation. The framework is evaluated using region-aware PSNR/LPIPS on edited and background pixels, alongside 3D reconstruction metrics against renderer-provided geometry.
Efficiency and Performance
JointEdit3D demonstrates strong efficiency, achieving inference times that are substantially faster than both feed-forward baselines under the full-3D protocol
by directly outputting RGB and geometry without an additional VGGT reconstruction pass. The framework shows competitive performance across various operations, with results on SceneEdit3D-Bench indicating superior edited-region quality and better background metrics compared to methods like Omni-3DEdit. The 3D geometry quality is also strong, achieving substantially better completeness, CD, and F-score
in the joint RGB + VGGT reconstruction variant. The ablation study confirms that components like the SceneAnchor Branch and edit-aware losses are crucial for achieving these gains.
Limitations
A key limitation noted is that JointEdit3D is reference-guided rather than a complete end-to-end text-to-3D editing pipeline,
meaning its quality depends on the edited reference image; errors in the reference can propagate.
Improvements for AI systems
Here are the specific improvements that an AI system, leveraging the JointEdit3D framework, could achieve, categorized by capability:
) 1. Enables High-Fidelity 3D Scene Editing via Latent Inpainting:
The system can perform scene editing directly within a unified RGB-geometry latent space rather than relying on separate 2D editing and 3D reconstruction pipelines. This means the AI can generate an edited scene where appearance synthesis (color, texture) and geometry (shape, structure) are jointly optimized from a single reference view.
) 2. Achieves Robust Edit Propagation Across Views:
By introducing the SceneAnchor Branch,
the system can propagate edit cues from a single edited reference frame across all other views in the source scene. This allows for geometrically coherent edits that maintain consistency across multiple perspectives, solving the problem of view inconsistencies common in decoupled methods.
) 3. Preserves Source Scene Structure Automatically (Mask-Free):
The AI can localize edits without requiring explicit edit masks during inference. The SceneAnchor Branch
acts as an implicit edit-region localizer by injecting source features as residual corrections, ensuring that unedited regions are preserved while the edited region departs from the source content.
) 4. Supports Complex, Multi-Operation Edits:
The system is capable of performing a wide range of scene modifications, including deletion (object removal), addition (insertion), relocation (moving objects to new positions), and appearance changes (material/texture replacement). The framework handles these complex tasks by learning edit-aware losses that prioritize modality-specific changes.
) 5. Provides Structure-Aware Geometric Supervision:
The training objectives are designed to be edit-aware,
decomposing the loss into separate terms for edited and background regions in both RGB and geometry latents. This ensures that the model learns to distinguish between what needs changing (the edit region) and what must remain stable (the background), leading to superior 3D structural completeness compared to standard MSE losses.
) 6. Optimizes for Sparse 3D Edits:
The system is specifically trained with a weighting scheme that upweights positions far from the reference frame and those exhibiting larger latent changes. This makes the framework highly effective for sparse edits, ensuring that small, localized changes (like removing a specific small object) are not diluted or overwhelmed by the preservation of large, unchanged background areas.
) 7. Delivers High-Quality 3D Geometry Reconstruction:
The system can reconstruct high-fidelity 3D geometry (point clouds and meshes) from the generated RGB-geometry latent, achieving competitive global metrics (CD, F-score) against established methods by jointly modeling appearance and structure.
In summary, the improved AI system can perform seamless, geometrically consistent 3D scene editing guided by a single reference image or video while preserving complex background structures across all views without needing manual edit masks during inference.
Sources
- Edit3r: Instant 3D Scene Editing from Sparse Unposed Images
- EditCast3D: Single-Frame-Guided 3D Editing with Video Propagation and View Selection
- The Replica Dataset: A Digital Replica of Indoor Spaces
- Wan: Open and Advanced Large-Scale Video Generative Models
- IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models