Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Rethinking Where to Edit".
Tom: Instruction-based image editing (IIE) models often suffer from over-editing, introducing unintended changes to regions unrelated to the desired edit.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: So, what’s the main thesis here regarding this paper, "Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing"? Basically, they are arguing that we need a better way to localize edits because current models do it blindly.
Jane: They claim that different editing operations—like adding something new versus removing something old—create distinct spatial patterns, and existing models don't account for this task dependency when localizing where the changes should happen.
Lu: Their key insight is that optimal localization is inherently task-dependent; for instance, object addition requires finding emergence in the target image stream, while removal focuses on content within the source image stream.
Meng: That sounds like a complex way to map instructions onto spatial regions, but I see how using intrinsic attention mechanisms could provide those initial coarse localization cues they mention.
Lalam: And they use those cues from the multi-modal joint attention module to derive an attention map, which is then aggregated across layers and tokens to create this "attention-derived mask Mimg(t)."
Tom: Right, so that's the first step: getting these initial spatial hints just by looking at how the model pays attention across different streams.
Jane: After getting those maps, they treat them as semantic indicators and then refine them in the latent feature space by extracting stream-specific latent features denoted as Fˆ(l)img (t).
Lu: The refinement step is where they get really smart; they build two feature centroids via masked average pooling, one specifically for edited regions and one for preserved regions.
Meng: So, instead of just using the raw attention map, they are creating two distinct semantic anchors—one for what's being changed and one for what needs to be kept—to make a decision.
Lalam: Then they assign every spatial token to the closest centroid based on cosine similarity to construct the final edit mask Mˆ(l)(t)i.
Conclusion: Tom: Looking at the title, "Rethinking Where to Edit," it really suggests a fundamental shift in how we approach instruction-based image editing by focusing on where the edits occur.
Jane: And the authors are basically proposing a training-free framework that leverages these stream-specific cues to guide latent updates during denoising, which is quite clever because it avoids needing extra training for localization.
Lu: The implications are big because if this works as claimed, we could move toward more semantically grounded editing where we don't introduce unintended changes outside the target area.
Meng: Practically speaking, that means less time spent on manual post-editing to clean up artifacts in areas the instruction didn't specify modifying.
Lalam: For culture and application, this advances AI by making it smarter about respecting content boundaries inherent in instructions, which is a big step for reliable content generation.
Tom: So, to sum up what we've heard about "Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing," it’s about using the model’s internal attention mechanisms to create task-aware masks that guide latent blending during denoising.
Jane: It simplifies the process by linking the required edit type—addition, removal, or replacement—to specific regions in either the source or target stream.
Lu: The long-term impact is showing that intrinsic model features can provide more reliable semantic cues than just raw attention activations for localization purposes.
Meng: I think it's robust because they showed the feature-based assignment is robust to the choice of attention threshold, which means we don't have to worry about tuning that parameter too much.
Lalam: It really shows how these architectures can be adapted to follow complex instructions with much higher fidelity in preserving what isn't being asked to change.
Jingxuan He, Xiyu Wang, Mengyu Zheng, Xiangyu Zeng, Yunke Wang, *Chang Xu
The University of Sydney
cs.CV
Submitted: 2026-04-22
Updated: 2026-09-28
Importance score: 81/100
The gist: Instruction-based image editing (IIE) models often suffer from over-editing, introducing unintended changes to regions unrelated to the desired edit.
Key concepts
- Task-Aware Edit Mask Construction
- The paper creates a specific mask for editing by looking at which image stream (source or target) is relevant for the task. For example, adding an object uses cues from the target stream, while removing one uses cues from the source stream. This ensures localization aligns with how different editing actions affect the images.
- Latent Feature Centroids
- The framework extracts latent features from each image stream and creates two 'centroids': one representing edited regions and another for preserved regions. By comparing a new token's feature to these centroids, the system can precisely decide whether a specific area should be edited or kept unchanged.
- Mask-Guided Latent Blending
- This technique controls how the image changes during the denoising process. It uses the task-aware mask to blend between the original target latent and a preserved latent version. This blending forces the update to only occur within the identified edit region, preventing unintended edits elsewhere.
Terminology
Summary
Instruction-based image editing (IIE) models often suffer from over-editing, introducing unintended changes to regions unrelated to the desired edit. This work proposes a training-free, task-aware edit localization framework that exploits intrinsic source and target image streams within IIE models to explicitly identify edit regions and guide latent updates, thereby enhancing content consistency without compromising instruction following.
How it works
The framework operates by decomposing and propagating attention activations from the multi-modal joint attention module to derive coarse localization cues. Specifically, it first derives cross-attention submatrices, such as the cross-attention submatrix A(l)img,ca (t) ∈ R img×Ntxt,
and self-attention submatrices (intra-modal interactions within each image stream
). These are combined through a diffusion of the linguistic signal to obtain an attention map, which is then aggregated across transformer layers and text tokens to yield an attention-derived mask Mimg(t).
How it works (Continued)
The paper further refines these initial cues by treating them as semantic indicators
and performing a subsequent refinement in the latent feature space. This involves extracting and normalizing stream-specific latent features, denoted as Fˆ(l)img (t),
and constructing two feature centroids via masked average pooling,
one for edited regions (C(l)img,1(t)) and one for preserved regions (C(l)img,0(t)). The final edit mask Mˆ(l)(t)[i] is then determined by assigning each spatial token i to the closest centroid based on cosine similarity.
The crucial innovation lies in the task-aware edit mask construction,
which selectively leverages source and target image streams based on the editing task. The paper identifies three representative subject-centric editing primitives:
-
Subject addition: localization is primarily manifested within the target stream, i.e.,
Mˆ (l) (t) = Mˆ(l)tgt (t).
-
Subject removal: regions are localized within the source stream, i.e.,
Mˆ (l) (t) = Mˆ(l)src (t).
-
Subject replacement: this task involves coupled semantic changes spanning both streams, i.e.,
Mˆ (l) (t) = Mˆ(l)tgt (t) ∪ Mˆ(l)src (t).
Once the task-aware edit mask is obtained, a mask-guided latent preservation
scheme is applied to constrain the evolution of target latents during denoising. This involves constructing an inverted latent Zinv (t)
defined as Zinv (t) = σtZtgt(0) + (1 − σt)Zsrc,
where σt is the noise schedule at timestep t. The target tokens are then updated via mask-guided latent blending
: Zˆtgt(t) = Mˆ(l)(t)Ztgt(t) + 1 − Mˆ(l)(t) Zinv (t).
This mechanism ensures that the evolution of target tokens is constrained to within the identified edit region, thereby enforcing localized edits while maintaining consistency in non-target regions.
The effectiveness of this approach is validated through systematic analysis on EdiVal-Bench. The study reveals that latent features provide more reliable semantic cues than attention activations,
as evidenced by higher IoU scores for feature-derived masks. Furthermore, the localization strategy is inherently task-dependent,
with semantic signals emerging from different image streams depending on the editing task type (e.g., addition semantics in the target stream vs. removal semantics in the source stream). The final results demonstrate that this method consistently improves content consistency in non-edited regions while maintaining strong instruction-following performance.
A systematic ablation study confirmed the robustness of the framework:
-
Attention threshold sensitivity: The feature-based assignment is
robust to the choice of attention threshold,
with similar IoU scores across a range of values. -
DiT layer depth: Segmentation performance improves monotonically as layer depth increases, peaking at layer 50, suggesting that
edit semantics are progressively formed and become more discriminative in deeper layers.
-
Denoising timestep: Applying latent preservation early (e.g., at t=5) significantly improves EdiVal-O, while applying it too late can lead to a
semantic mismatch,
indicating that the choice of timesteps must be balanced for optimal fidelity and consistency.
In summary, the proposed framework addresses over-editing by explicitly identifying edit regions through task-aware mask construction based on stream-specific semantic cues derived from latent features. This localization is then enforced via mask-guided latent blending, leading to "semantically grounded editing with minimal unintended changes.
Improvements for AI systems
Based on the provided scientific paper, here are specific, actionable improvements for existing AI image editing systems and what those improved systems could achieve:
The core improvement is shifting from task-agnostic
editing to a task-aware localization
mechanism that explicitly identifies and isolates the edit region based on the nature of the instruction.
Here are the specific enhancements:
Improving Non-Edit Region Consistency (Reducing Over-Editing):
By implementing a training-free, task-aware edit localization framework (using attention cues to derive feature centroids and task-dependent mask construction), existing models can be guided to preserve non-target regions with significantly higher fidelity. The system will stop modifying background elements, human features, or unrelated objects when the instruction only targets a specific object or region.
Task-Specific Edit Localization:
The system will dynamically adjust its localization strategy based on the editing primitive:
-
For Subject Addition: It will prioritize localization signals from the target image stream to ensure new content is placed contextually within the scene, rather than drifting into unrelated parts of the source image.
-
For Subject Removal: It will focus localization signals on the source image stream to ensure only existing content is erased, preventing accidental corruption of adjacent background elements.
-
For Subject Replacement: It will jointly leverage cues from both streams (source for removal and target for generation) to achieve a precise
cut and paste
effect without introducing artifacts at the seam.
Superior Semantic Grounding via Latent Features:
The improved system will utilize latent features derived from the joint attention module as primary semantic indicators, rather than relying solely on attention maps. This leads to more robust and semantically coherent partitioning of tokens, resulting in cleaner boundaries and more complete spatial coverage for the edit region compared to methods using only attention activations.
Enforced Spatial Constraint during Denoising (Mask-Guided Latent Preservation):
The system will introduce a mechanism that selectively constrains the evolution of target image tokens during denoising. By blending between the original source latent and the modified target latent only within the task-aware edit mask, it prevents all-to-all coupling
from treating all spatial tokens as equally mutable, thereby enforcing localized updates and dramatically reducing unintended changes.
Improved Handling of Complex Editing Scenarios:
The proposed framework will systematically address complex editing needs by:
-
Using the union of source and target masks (for replacement) to ensure both removal and addition are perfectly synchronized.
-
Implementing post-processing steps (connected component retention, hole filling, spatial expansion) to handle disjoint edit regions more effectively than current methods, ensuring full coverage even when multiple distinct edits are requested simultaneously.
The resulting improved AI system can perform the following:
Generate images with significantly higher fidelity to the intended edit instructions while maintaining near-perfect preservation of the original image's non-edited content (e.g., high-quality subject replacement without object bleeding
or background alteration).
-
Execute complex, multi-step edits (like replacing a specific object) with precise spatial control over where the modification occurs, minimizing visual artifacts at boundaries.
-
Provide superior performance on specialized benchmarks like EdiVal-Bench, demonstrating enhanced content consistency and instruction following across various editing primitives compared to current state-of-the-art models (e.g., Step1X-Edit or Qwen-Image-Edit).
Sources
- SAM 3: Segment Anything with Concepts
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- DiffEdit: Diffusion-based semantic image editing with mask guidance
- Guiding Instruction-based Image Editing via Multimodal Large Language Models
- Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive Activations
- ACE: All-round Creator and Editor Following Instructions via Diffusion Transformer
- Prompt-to-Prompt Image Editing with Cross Attention Control
- Direct Inversion: Boosting Diffusion-based Editing with 3 Lines of Code
- Flow Matching for Generative Modeling
- Step1X-Edit: A Practical Framework for General Image Editing
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
- Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention
- InstructEdit: Improving Automatic Masks for Diffusion-based Image Editing With User Instructions
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Qwen-Image Technical Report
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Qwen2.5 Technical Report
- Group Relative Attention Guidance for Image Editing
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models