DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes".
Jane: The gist: DORS proposes a training-free,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, looking at the paper DORS: Dynamic Attention Routing for Diffusion-based Object Removal in Dense Scenes by Haitong Tang, Haipeng Liu, and Yang Wang... it’s a training-free framework that uses dynamic attention routing to manage information flow during diffusion.
Jane: The authors show that this approach, combining Instance-Filtered Attention and Context-Guided Routing, effectively suppresses instance interference while restoring necessary structural cues in dense scenes.
Lu: It gives us a way to understand exactly how misleading semantic information propagates through the attention space when dealing with similar objects.
Meng: From an engineering standpoint, the plug-and-play nature is valuable because it means we can apply this mechanism without needing extensive retraining or complex conditional guidance setup.
Tom: That’s right; it’s a direct solution to the instability of object removal in crowded environments, showing strong results like a ReMOVE score of seventy-five point five three on DOR-Bench.
Jane: The authors acknowledge that while this is effective at removing the target instance, it stops short when dealing with effects like shadows and reflections that stay behind.
Lu: They suggest that future work should involve jointly identifying and removing those visual effects along with the objects themselves to improve realism.
Meng: So, DORS provides a solid method for scene understanding in AI manipulation right now, even if it doesn't solve the problem of perfect visual realism yet.
Conclusion: Tom: So, DORS is about using dynamic routing in attention to remove objects without needing any training data for this specific task.
Jane: It’s essentially a way to tell the AI where to look so it doesn't get confused by too many similar things when removing something dense.
Lu: The core idea is that you can control the flow of information within the attention space instead of just letting it happen naturally.
Meng: So, if I understand this right, they’re not retraining the whole model to handle crowded scenes; they're just tweaking how the existing network pays attention during image editing.
Lalam: From a cultural perspective, this means we can make AI tools that work reliably in real-world messy situations much faster and more consistently.
Tom: Exactly. The authors call it Instance-Filtered Attention and Context-Guided Routing, which sounds complicated, but the goal is simple: stop the AI from mixing up similar objects.
Jane: It’s about shielding the removal process so it only focuses on what's actually there and ignores the noise from nearby instances.
Lu: They show that this mechanism is effective because it doesn't just block everything; it dynamically adjusts how much focus goes to different parts of the scene based on density and distance.
Meng: I see a practical application here where you could have more stable results in industrial quality control images, for example, where objects are packed closely together.
Lalam: And if we can achieve that level of consistency without constant retraining, it opens up a lot of possibilities for deploying these kinds of tools everywhere.
Tom: Right. So DORS shows that by modeling the information flow explicitly during denoising, we get much cleaner object removal in those tricky dense scenarios.
Jane: It’s a really neat piece of research because it gives us a principled way to tackle the problem of instance interference head-on.
Lu: The limitation they mention is that it still struggles with visual effects attached to the object, like shadows or reflections, which is something we definitely need to look at next.
Meng: That makes sense; if you remove an object but leave behind a shadow that doesn't match the background, the result isn't complete yet.
Lalam: So the next step for this kind of work is figuring out how to get the AI to handle those complex visual interactions too.
School of Computer Science and Information Engineering, University of Technology
cs.CV
Submitted: 2026-07-18
Updated: 2026-10-08
Code: https://github.com/httang1224/DORS
Importance score: 92/100
The gist: The gist: DORS proposes a training-free, plug-and-play framework that formulates object removal as semantic information flow control in the attention space to effectively suppress misleading
Key concepts
- Instance Interference
- This happens when the model's attention mechanism incorrectly focuses on semantically similar objects in a dense scene instead of the target object. This causes erroneous information propagation, leading to unstable and inaccurate object removal because the model gets confused by redundant visual cues.
- Instance-Filtered Attention (IFA)
- IFA acts as a semantic shield. It uses instance segmentation to identify similar objects and then dynamically prunes attention connections from the masked area to these similar regions. This prevents misleading information from spreading across similar instances, effectively blocking unwanted cross-region interactions.
- Context-Guided Routing (CGR)
- CGR is a dynamic mechanism that decides how much attention to give the full context versus the filtered instance information. It uses a spatially adaptive weight based on scene density and distance to similar objects. This ensures that useful background structure is preserved in close areas while suppressing interference in distant regions.
Terminology
Summary
The gist: DORS proposes a training-free, plug-and-play framework that formulates object removal as semantic information flow control in the attention space to effectively suppress misleading information from similar instances while preserving structural consistency in dense scenes.
Problem and Motivation
Existing methods often fail to completely erase target instances and leave residual artifacts, especially in dense scenes containing multiple similar instances, because they rely on unconstrained context aggregation which leads to Instance Interference
where masked queries tend to align with semantically similar keys due to global similarity matching in self-attention. This issue arises from erroneous information propagation in the attention space, resulting in uncontrolled cross-region information aggregation and erroneous feature propagation. Consequently, mask constraints or conditional guidance alone are often insufficient to reliably prevent such misallocation of attention in dense scenes, leading to unstable and inaccurate object removal.
Proposed Framework (DORS)
DORS introduces a dynamic attention routing mechanism built upon two complementary components: Instance-Filtered Attention (IFA) and Context-Guided Routing (CGR). This mechanism explicitly regulates cross-region information flow within the attention space. IFA acts as a semantic shield by dynamically constructing constraints to prune misleading attention connections to semantically similar but irrelevant regions. CGR dynamically redirects attention toward valid contextual cues, thereby preserving structural consistency. Together, this complementary “shield-and-bridge” design suppresses interference while retaining meaningful background information.
Dynamic Attention Routing Mechanism
The framework operates by fusing two complementary pathways: a full-context branch (OF) derived from the mask-based bias (which suppresses intramask propagation) and an instance-filtered branch (OP) which blocks interactions with target-similar regions. A dynamic, spatially adaptive routing mechanism fuses these two branches according to the global distribution and local geometric relationships of target-similar instances. This is achieved by defining a spatial adaptive filtering weight w(x) that jointly accounts for global instance density and local distance to target-similar instances. The final attention output is computed as O = w(x) Op + (1 - w(x)) Of, resulting in a distribution- and geometry-aware routing mechanism.
Stage-wise Application and Control
The proposed attention control is primarily applied during the early denoising stages, when the global structure is established, and is relaxed in later stages to facilitate detail refinement. This stage-wise design effectively balances structural consistency and fine-grained detail generation. The dynamic weight w(x) incorporates a global density measure ρ = area(Ms) / area(unmasked region), where a larger ρ indicates stronger semantic interference, necessitating more aggressive suppression. Furthermore, local spatial relationships are modeled by computing the distance di(x) to nearby similar instances. The adaptive filtering weight w i(x) is defined based on these principles, assigning lower filtering weights to proximal regions to preserve structural continuity and progressively increasing them for distant regions where contextual information is more likely to introduce semantic interference.
Evaluation and Results
Extensive experiments demonstrate that DORS outperforms state-of-the-art methods, particularly in reducing incomplete removal and duplicate artifacts. On the DOR-Bench benchmark, DORS achieves a ReMOVE score of 75.53, significantly reducing MSN to 1.25 and MARS to 0.70 compared to the second-best results. The method also outperforms ObjectClear in background fidelity, improving PSNR from 29.83 to 35.06 while reducing LPIPS from 7.43 to 3.14. Qualitative analysis confirms that DORS achieves more complete removal with fewer artifacts while preserving structural and textural consistency across diverse scenes. The method is also robust across different backbones, consistently improving performance on SD1.5, SD2.0, and SDXL models.
Conclusion
DORS successfully addresses object removal in dense scenes by explicitly modeling information flow in the attention space and introducing dynamic attention routing to selectively regulate cross-region interactions. The combination of Instance-Filtered Attention and Context-Guided Routing enables fine-grained control over attention information, resulting in stable and visually consistent removal results. The framework is a training-free, plug-and-play solution that demonstrates strong performance across various benchmarks and backbones. The design validates the effectiveness of each component in DAR, confirming that IFA suppresses misleading semantic information while CGR restores useful structural cues. This approach provides a principled explanation for Instance Interference in dense scenes. The paper concludes by noting that DORS remains limited in handling complex object-associated visual effects beyond the target appearance itself, such as shadows and reflections. The research suggests that jointly identifying and removing target objects together with their associated visual effects remains an important direction for improving the completeness and realism of object removal in complex scenes. The overall design confirms that IFA and CGR play complementary roles in suppressing interference while restoring valid contextual information. The study concludes by stating that DORS effectively mitigates semantic interference from similar instances while preserving structural continuity.
How it works
The framework is built upon a diffusion-based image inpainting process where the noisy latent zt is updated using a modified noise prediction network. The core modification involves replacing the standard self-attention computation with the fused attention mechanism derived from IFA and CGR. This fusion occurs during early denoising stages to establish global structure, and later stages revert to the standard predictor for detail refinement.
Instance-Filtered Attention (IFA)
IFA performs semantic pathway pruning by blocking attention from the masked region to both itself and target-similar regions in the surrounding area. This is achieved by first performing instance-level segmentation using SAM3 model to identify similar instances, denoted as Ms, and then introducing a similarity-aware pruned Density- and distance-conditioned Piecewise Mapping f d!(σ, ρ). The resulting attention output Op is computed as softmax(S + Bp) V, where Bp is the bias matrix that dynamically prunes attention pathways from masked queries to both intra-mask and target-similar regions.
Context-Guided Routing (CGR)
CGR selectively restores useful contextual information by fusing the full-context branch (Of) with the instance-filtered branch (Op) based on a dynamic, spatially adaptive routing mechanism. The routing weight w(x) is defined as max i w i(x), where w i(x) is determined by a piecewise function that jointly accounts for global instance density and local distance to target-similar instances. This dynamic weighting ensures that the full-context pathway is favored in proximal regions to maintain structural continuity, whereas the filtered pathway dominates distant regions to suppress misleading semantic propagation.
Design Ablation
A comprehensive ablation study confirms the necessity of both components; adding only the standard full-context branch (FCB) provides limited improvement, whereas IFA substantially reduces MSN and MARS to 11.75 and 5.43, respectively. Building upon IFA, CGR further improves the scores to 1.25 and 0.70 with only a marginal increase in inference time. The sensitivity analysis of the density control parameter α shows that the best performance is achieved at α = 0.5, where the model achieves a good balance between interference suppression and structural preservation.
Effect of Segmentation Models
The robustness of DORS to imperfect mask inputs is demonstrated across different segmentation models; DORS (SAM3) achieves MSN=4.17 and MARS=2.38, while the Oracle masks yield MSN=2.08 and MARS=1.99. Higher-quality masks generally lead to better removal performance, although the performance gap between SAM3 and SAM1 remains limited. The paper concludes by noting that even with the weakest SAM1-based variant, DORS still achieves strong removal performance and outperforms existing state-of-the-art methods.
Failure Cases
DORS remains limited in handling complex object-associated visual effects beyond the target appearance itself, such as shadows and reflections. For example, a shadow cast by the removed object may persist on the background surface, reducing both visual realism and overall consistency. This limitation stems from the primary objective of DORS, which is to suppress misleading semantic information from similar instances while preserving useful contextual cues from unmasked regions for background reconstruction. The study concludes by stating that jointly identifying and removing target objects together with their associated visual effects remains an important direction for improving the completeness and realism of object removal in complex scenes.
Improvements for AI systems
- Bold header: Dynamic Attention Routing Mechanism Implementation
The system can now explicitly regulate information flow by integrating Instance-Filtered Attention (IFA) to suppress misleading semantic information from similar instances
and Context-Guided Routing (CGR) to dynamically route complementary scene information to maintain visual consistency.
- Bold header: Stage-wise Control Strategy
The model gains a structured denoising process where it can apply control only during critical stages, as the paper states, the proposed attention control is primarily applied during the early denoising stages, when the global structure is established, and is relaxed in later stages to facilitate detail refinement.
- Bold header: Spatially Adaptive Routing Weighting
The AI can dynamically adjust its attention based on both global density and local geometry using a formula that combines these factors, as described by the weighting function w(x) = max i w i(x)
which balances structural continuity against semantic interference.
- Bold header: Robustness to Imperfect Segmentation
The system can maintain high performance even when input masks are imperfect, as demonstrated by the finding that even with the relatively weaker SAM1 [17], DORS still achieves strong removal performance and outperforms ObjectClear and AttentiveEraser.
- Bold header: Systemic Benchmark for Dense Scenes
The framework provides a rigorous evaluation tool through DOR-Bench, which specifically targets scenarios with strong semantic ambiguity caused by similar instances,
allowing developers to systematically test and improve models in challenging environments.
- Bold header: Human and VLM Consensus Alignment
The AI's output quality can be validated against human experts and multimodal models, achieving a Spearman rank correlation coefficient of 0.90 over five methods,
ensuring that the removal accuracy is robust across different evaluation paradigms.
Sources
- SAM 3: Segment Anything with Concepts
- Classifier-Free Diffusion Guidance
- Auto-Encoding Variational Bayes
- Flow Matching for Generative Modeling
- Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Selective Coupling of Decoupled Informative Regions: Masked Attention Alignment for Data-Free Quantization of Vision Transformers
- Denoising Diffusion Implicit Models
- OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
- Precise Object and Effect Removal with Adaptive Target-Aware Attention
- GeoRemover: Removing Objects and Their Causal Visual Artifacts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models