To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion".
Jane: The paper was written by Shaswati Saha, Rajasekhar Anguluri and Manas Gaur from University of Maryland Baltimore County, USA.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary and Core Mechanism: Tom: The core of the issue addressed in "To Erase, or Not to Erase" is how traditional concept erasure methods—CETs—fail because they rely on static lists or proxy embeddings. They struggle to capture the dynamic way a prompt steers a model.
Jane: So, this new approach uses what they call "Knowledge Search" to fix that. Instead of relying on predefined banks, the system treats the target concept like a query and searches through the entire model vocabulary for relevant tokens.
Tom: This is done in two stages: first to coarsely rank candidates based on how similar they steer the denoising process, and then they use a much finer filter.
Lu: That’s exactly where the power lies; it allows us to capture those subtle connections between concepts that traditional methods miss because it looks directly at the generative geometry of finding related tokens.
Meng: I am curious about this two-stage filtering mechanism—is it just selecting similar things, or is there a complex process ensuring we are only getting exactly what's needed?
Lalam: The second stage ensures that we are keeping only those concepts that are truly tied to the target in the images. This keeps our output highly relevant to what we intend to remove while maintaining visual fidelity and relevance.
Tom: Once those lists of things to erase and things to keep, they apply a preservation-aware subspace projection. This is the crucial mathematical edit that allows everything else stay exactly where it belongs in the model space.
Jane: It’s not just deleting target directions; they are specifically protecting the directions that represent benign concepts, so we don't accidentally ruin other parts of the model or suppress unrelated ideas during erasure.
Tom: This approach appears to avoid all that destructive over-suppression that was a major flaw in older methods. It is about surgical precision, removing only what is required.
Lu: I really appreciate that they are orthogonalizing the space, which creates a clean, mathematically defined break between the erased concepts and the retained ones, making them distinct conceptually.
Meng: The way they apply this to latent diffusion models suggests a very strong potential for applying this technique to various generative AI systems across many different platforms.
Lalam: Lalam sees this as a path that allows us to uphold safety standards while preserving the high quality and aesthetic value of the images we are generating.
Improvements and Robustness: Tom: The authors identify three key enhancements in "To Erase, or Not to Erase": dynamic knowledge search, preservation-aware projection, and a new adaptive mechanism called Adaptive Subspace Expansion or ASE.
Jane: The entire premise of this ASE is that even after we successfully erase a concept, an attacker might find subtle ways to get it back using paraphrasing or similar prompts. This technique addresses that future vulnerability directly.
Tom: They use iterative searches to discover these "re-emergence triggers"—things like textual inversions—and then they adaptively expand the erased subspace to cover those triggers too.
Lu: This iterative process is genius because it is not a one-time fix; it is actively finding and patching potential loopholes in the model's understanding of concepts, which shows a deep level of foresight into adversarial thinking.
Meng: This adaptive expansion suggests they are making the model much more resilient to unexpected input variability, which is critical for deployment in real-world systems where inputs are always changing.
Lalam: Lalam thinks the ability to handle these potential triggers without destroying other concepts is vital for cultural integrity. If we cannot re-emerge unwanted imagery under attack, our AI can be trusted by users more readily to produce appropriate content.
Tom: That leads us into how they measure success with a new metric called BEUS, or Balanced Erasure Utility Score.
Jane: It’s not just measuring how well the erasure works; it measures the careful trade-off between that erasure and the model's original utility for benign concepts.
Tom: The authors use a harmonic mean aggregation in BEUS, which is a very thoughtful way to show that if one metric starts getting worse, both metrics are penalized.
Lu: That mathematical structure of the harmonic mean shows how severe the penalty is when it is really failing to balance, which is incredibly powerful for safety-critical applications.
Meng: Practically, this means we get a holistic view of the system's quality and its overall reliability rather than just looking at a single percentage of success or failure.
Lalam: Lalam appreciates how this metric allows for cultural compliance; it quantifies precisely how much safety comes at the cost of usability in real-world AI implementation.
Conclusion and Wrap-up: Tom: We have seen how "To Erase, or Not to Erase" uses a dynamic approach to solve the old problem of static concept banks, which is a significant leap forward in making AI safer and more predictable.
Jane: The core message is that we can achieve strong robustness against re-emergence without sacrificing the model's ability to generate high-quality, non-target content.
Tom: As we wrap up this discussion, let's make sure we give one last thought on the impact of each of us.
Lu: I think we should all recognize how much more complex AI is becoming; this shows that even our understanding of its behavior needs to evolve to keep pace with these sophisticated tools.
Meng: From a deployment standpoint, it is encouraging to see a solution that is both robust and practical for real-world integration into production systems.
Lalam: Lalam believes this paper opens doors for better cultural stewardship, ensuring that the future of generative AI is both powerful and responsible in its use.
Tom: That's a wonderful final thought, Jane; thank you all for sharing your insights on "To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion."
Final Wrap-up: Tom: We have so much more to discuss about the future of AI safety and adaptability regarding how we can make these tools even more robust.
Jane: But before we go, I want to hear one last quick reaction from Lu, Meng, and Lalam.
Lu: I am really excited to see how this knowledge-based approach scales up when multiple concepts are involved in the same prompt.
Meng: The efficiency of the core algorithm is certainly something that warrants further investigation into real-world operational costs for me.
Lalam: I hope this enables a future where we can trust generative AI to create images that are both beautiful and appropriate for cultural consumption.
Tom: That is a wonderful final thought, Jane; we will see you next time with another exciting paper!
University of Maryland Baltimore County, USA · University of Maryland Baltimore County, USA
cs.CV, cs.LG
Submitted: 2026-07-26
Updated: 2026-09-04
Comments: Accepted to ECCV 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 91/100
The gist: * Concept Erasure Techniques (CETs) aim to suppress user-specified targets (e.g., NSFW content or copyrighted styles) in text-to-image diffusion models while preserving model utility for benign
Key concepts
- Knowledge Search
- Instead of relying on predefined banks, this system treats the target concept like a query and searches through the entire model vocabulary for relevant tokens. This allows it to capture subtle connections that traditional methods miss by looking directly at the generative geometry of finding related tokens.
- Preservation-aware Subspace Projection
- This is a crucial mathematical edit applied after identifying what to erase and keep. It ensures that only target directions are removed, specifically protecting the directions representing benign concepts so other parts of the model are not accidentally ruined or suppressed during erasure.
- Adaptive Subspace Expansion (ASE)
- This mechanism addresses future vulnerabilities where an attacker might re-emerge a concept using paraphrasing. The system uses iterative searches to find these "re-emergence triggers" and then adaptively expands the erased subspace to cover those potential loopholes.
- BEUS (Balanced Erasure Utility Score)
- This is a new metric used to measure success. It quantifies the careful trade-off between how well the erasure works and the model's original utility for benign concepts, using a harmonic mean aggregation to penalize poor balance.
Terminology
Summary
Concept Erasure Techniques (CETs) aim to suppress user-specified targets (e.g., NSFW content or copyrighted styles) in text-to-image diffusion models while preserving model utility for benign concepts. However, existing CETs suffer from a critical trade-off: robust erasure often degrades model utility, and vice versa.
A key weakness in current methods is the definition of what to erase and what to preserve. Many rely on static concept banks (manually specified or retrieved using proxy CLIP embeddings). These static banks fail to model how prompts steer the denoising trajectory in diffusion models, leaving edited models vulnerable to re-emergence attacks. Furthermore, adversarial robustness often leads to over-suppression of semantically adjacent concepts, compromising generation quality.
We introduce Preservation aware Adaptive Ranked Subspace Expansion (PARSE), a training-free framework designed for robust concept erasure in latent diffusion models. Unlike existing methods, PARSE treats the definition of erase and retain sets as a diffusion-grounded retrieval problem,
utilizing the model's inherent knowledge rather than static anchors.
PARSE operates through three integrated components:
1. Diffusion–ground Knowledge Search (Knowledge Bank Retrieval):
Instead of relying on proxy embeddings, PARSE dynamically discovers target-adjacent concepts by querying the diffusion model using classifier-free guidance (CFG). This process constructs an erase bank (E) and a retain bank (R).
-
Stage 1: Coarse Ranking: The system computes the CFG steering embedding S(v) for every token v in the frozen text-encoder vocabulary V (Eq. 5). Tokens are ranked by their cosine similarity to the target concept's steering embedding, cosine(v, target) (Eq. 6), yielding a set of top candidates (V top-K').
-
Stage 2: Fine-grained Filtering: To eliminate false positives (e.g, retaining
fruit
when erasingapple
), the Stage 1 candidates are filtered using CLIP classification on cached images I(v) i. The the final erase token bank E is defined as those tokens whose cached images are consistently classified as containing the target (Eq. 8).
2. Preservation-aware Subspace Projection:
The erased model M target is created by applying a closed-form linear edit to the cross-attention value space: V edited = P V prompt. The preservation-aware projection ensures that the erase basis E is orthogonal to the retain basis R:
P = I - E E, where E's columns are re-orthonormalized after applying (I - RR)
This projection explicitly protects retain semantics while removing target directions (Eq. 10).
3. Adaptive Subspace Expansion:
To address triggers that lie outside the initial vocabulary search space, PARSE iteratively searches for re-emergence triggers using Textual Inversion. It then adaptively expands the erased subspace by appending new trigger directions (enew) to the erase basis, provided that this expansion does not conflict with retain semantics (i.e., maintaining span(E) span(R)). This iterative process refines the projector P (Algorithm 1).
To quantify the balance between erasure robustness and utility, we introduce the Balanced Erasure Utility Score (BEUS).
Balanced Erasure Utility Score (BEUS):
BEUS is defined as a harmonic-mean aggregation of scaled Attack Success Rate (ASR sc) and scaled Fréchet Inception Distance (FID sc):
BEUS = 2 times ASR sc times FID sc / (ASR sc + FID sc)
This metric is bounded in [0, 1] and is strictly increasing in both ASR sc and FID sc. BEUS is high only when the model simultaneously achieves strong erasure robustness (low ASR) and preserves utility (low FID).
Experiments across three categories—NSFW (Nudity), Style (Van Gogh), and Object (Garbage Truck)—demonstrate that PARSE achieves state-of-the-art performance:
-
Erasure Effectiveness: PARSE achieves extremely low target ASR values, such as 2.09 for NSFW and 0.00 for style and object categories (Table 1).
-
Robustness: The method is the most robust across all concept categories, consistently preventing re-emergence under various attacks (CCE, UD, RAB), unlike many baselines.
-
Utility Preservation: PARSE preserves near-perfect utility. Unlike other methods that over-suppress or degrade adjacent concepts (e.g., failing to generate a
tow truck
after erasing agarbage truck
), PARSE maintains the ability to generate these benign, semantically related concepts while blocking target re-emergence (Fig. 4). -
Analysis: Further analysis confirms that the diffusion-grounded Knowledge Search outperforms proxy VLM approaches, and that PARSE is robust to various hyperparameter choices (K', eta, s).
Improvements for AI systems
The implementation of Preservation-aware Adaptive Ranked Subspace Expansion (PARSE) introduces several critical enhancements to existing AI systems that utilize latent diffusion models (Text-to-Image generation). These improvements move beyond static, proxy-based concept erasure, addressing fundamental flaws in robustness and utility.
Implementation: The system replaces pre-defined or LLM-curated concept banks with a dynamic, diffusion-grounded retrieval mechanism (Knowledge Search). It queries the frozen text encoder vocabulary (V) by calculating the Classifier-Free Guidance (CFG) steering embedding (epsilon = epsilon c - epsilon u) for target concepts.
Improvement: The system identifies target-adjacent concepts based on their actual influence on the denoising trajectory within the latent space, not just their similarity in a static CLIP embedding space. This ensures the erasure bank (E) is highly relevant to the model’s internal generative geometry, making it significantly more effective than previous methods.
Implementation: The system utilizes a specialized closed-form projection (P = I - E E) that is explicitly designed to be retrain-free. Crucially, the projection is preservation-aware, ensuring that the target subspace (span(E)) is orthogonalized against the retain subspace (span(R)).
Improvement: This directly solves the over-suppression
problem. The system can now erase a target concept (e.g., garbage truck
) without inadvertently corrupting semantically adjacent or benign concepts (e.g., tow truck
or trash can
), thereby preserving high model utility on non-target prompts.
Implementation: Instead of performing a single, fixed edit, the system iteratively searches for re-emergence triggers using Textual Inversion (TI). The erasure subspace is only expanded (E to E new) when a newly discovered trigger direction conflicts with the established retain semantics (span(R)).
Improvement: This provides dynamic, adaptive robustness. The system continuously hardens itself against subtle, adversarial, or paraphrased prompts that attempt to re-emerge the erased concept.
Implementation: The system incorporates the Balanced Erasure Utility Score (BEUS) as a primary performance metric, aggregating Attack Success Rate (ASR) and Fréchet Inception Distance (FID) via a harmonic mean.
Improvement: This allows for rigorous, quantitative validation of achieving the intended balance between erasure efficacy and utility preservation—a critical feature that is typically assessed upon two separate metrics in existing systems.
The improved AI system, utilizing PARSE architecture, can:
-
Achieve State-of-the-Art Robustness: It will reliably suppress target concepts (e.g., NSFW content or specific object types) even when presented with sophisticated adversarial attacks, paraphrased prompts, or attempts to bypass the edit using semantic neighbors.
-
Maintain High Generative Utility: It will generate high-quality images for benign/non-target prompts that are comparable to the original unedited model, avoiding the
over-suppression
and degradation issues common in older methods (e.g., successfully generating aperson
after erasing nudity). -
Scale Concept Erasure Dynamically: It can manage complex, multi-concept erasure tasks by treating each target as a parallel query against its own dynamic knowledge bank, ensuring consistency across diverse concept categories (e.g., erasing multiple artists or various types of vehicles).
Sources
- Qwen3-VL Technical Report
- On Evaluating Adversarial Robustness
- Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
- TraSCE: Trajectory Steering for Concept Erasure
- SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
- Understanding Diffusion Models: A Unified Perspective
- Circumventing Concept Erasure Methods For Text-to-Image Generative Models
- SeeBel: Seeing is Believing
- Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
- EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
- Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models