Erased but Not Forgotten: How Backdoors Compromise Concept Erasure
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: I'm Nadia, and with me are Elias and Priya, guest researcher.
Elias: Today's paper: "Erased but Not Forgotten".
Nadia: A critical vulnerability has been uncovered where deliberately inserted triggers can survive concept erasure techniques in text-to-image diffusion models, posing a significant risk to content safety efforts.
Elias: First, who's behind it and why it matters.
Title and authors: Nadia: So, we're looking at the paper "Erased but Not Forgotten: How Backdoors Compromise Concept Erasure," which really digs into how deliberately inserted triggers can bypass concept removal techniques in text-to-image diffusion models. It points out that even methods designed to erase concepts might fail if a malicious link is established first.
Elias: I agree, Nadia, the title itself sets the stage for something quite alarming: it suggests that what we think is a clean erasure process could be fundamentally flawed because of these hidden triggers. The authors are showing that existing erasure methods don't always sever those connections cleanly.
Priya: From a measurement standpoint, this is significant because it moves the focus from just checking if an output looks right to understanding the underlying mechanism and whether that connection actually stays or disappears after training modifications.
Nadia: Exactly, Priya; the core of what they're showing is that this Erasure Evasion Backdoor, or EEB, can survive even quite robust erasure processes. They demonstrate that these backdoors persist across various erasure methods, which is a very worrying finding for content safety efforts.
Elias: The paper details three ways this threat model can manifest: in the black-box setting through simple data poisoning, and in the white-box setting with variations like EEBsurface, EEBshallow, and EEBdeep. This shows that the vulnerability isn't limited to just one way of training or access control.
Nadia: And what they really highlight is that their goal for an adversary is twofold: to embed triggers that keep access to the target concept after erasure, while still making the poisoned model look identical to a clean one when prompted with normal inputs. That's a very tricky thing for anyone trying to secure these models.
Priya: From my perspective, seeing this across black-box and white-box scenarios means that the risk isn't confined to just being able to tamper with training data or just having direct access to the weights; it suggests a systemic fragility in how we define and implement concept removal.
Elias: Indeed, Priya, and the paper lays out specific mechanisms for realizing these attacks: EEBdata through dirty-label poisoning, RICKROLLING which modifies the text encoder by minimizing cosine similarity between trigger and target embeddings, EVILEDIT that alters cross-attention key/value mappings using a closed-form solution to align the trigger with the target.
Nadia: Those specific attack mechanisms are what make this paper so concrete; we're not just talking about abstract risks, we're seeing exactly how an adversary can implement these links across different parts of the AI system.
Title and authors: Priya: The authors then introduce EEBdeep as a score-based attack that injects the trigger across the entire diffusion pipeline by optimizing an objective function that balances trigger loss, retention loss, and quality loss. This deep intervention is what seems to be where the persistence really becomes substantial.
Elias: That score-based approach is interesting because it aims for a holistic injection, but the results suggest that EEBdeep remains effective across all six state-of-the-art erasure methods they tested. They found it was generally the most persistent, with example figures showing up to eighty-two percent success against celebrity identity unlearning <ref:2504.21072#pg0,up to 82% success against celebrity identity unlearning>.
Nadia: That persistence is what really strikes me; they show that even when you use methods specifically designed to find alternative representations of the target concept, EEB deep keeps it alive at a high rate. I'm also paying attention to the comparison point: for celebrity identities, EEBdeep generated the target concept in seventy-nine point seven two percent when prompted with the trigger, compared to only eight point seven six percent when conditioned on the target concept itself under RECE erasure.
Priya: That quantitative difference suggests that relying solely on iterative counterfactual training loops or other methods for robustness isn't enough to guarantee complete removal; there's a gap that these targeted backdoors exploit. I see this as a strong indication that we need more rigorous testing protocols before deploying any concept removal tool.
Elias: And the paper doesn't just stop at showing what works; they also provide diagnostic utility for stress-testing future erasure techniques, suggesting researchers can use these controlled backdoors to figure out which methods are achieving true semantic removal versus just superficial concealment.
Nadia: That diagnostic potential is huge for us as security researchers because it gives us a way to actively probe the limits of our defenses rather than just assuming a method works. It helps us distinguish genuine unlearning from something that only hides access paths temporarily.
Priya: And thinking about the implications for privacy, if these backdoors can survive erasure, it means that sensitive concepts like celebrity identities or specific objects could resurface unexpectedly, even after we've tried to sanitize the model. This raises serious questions about the long-term reliability of content filtering systems.
Elias: From a cryptographic viewpoint, this vulnerability highlights how easily localized modifications in embeddings or cross-attention layers can be leveraged to maintain a specific relationship that was supposed to be severed by training procedures like those used in UCE or MACE.
Title and authors: Nadia: So, when we look at the improvements suggested by the authors, they are essentially calling for an adversarial stress-testing protocol where any new erasure method gets tested against EEBdeep first to see if it can actually sever that deep link. That sounds like a necessary step for anyone developing these tools.
Priya: I also think their suggestion about adaptive erasure strategies based on the architectural layer where the backdoor is most likely to persist is important, because it points toward a more granular defense mechanism rather than just treating all models and all attacks with the same approach.
Elias: That dynamic unlearning system idea makes sense because EEB surface targeting the text encoder behaves differently than EEBdeep targeting the U-Net backbone; focusing defenses where the persistence is highest seems like an efficient use of resources.
Nadia: And then there's their proposal for enhanced detection mechanisms using inference-time activation monitoring, which could flag triggered prompt patterns or anomalous activations correlated with known triggers, especially after a model has been sanitized using EEBdeep. That offers a way to catch the output before it even reaches the user.
Priya: Real-time monitoring is certainly appealing because it provides a continuous defense layer during inference, which is vital for catching things that might slip past static training-based erasures and ensuring that the model behaves as intended in production environments.
Elias: They also discussed optimizing hyperparameters, showing that setting the interpolation weight alpha to zero point five offers a good balance between achieving high trigger accuracy, around ninety-three point four percent, while still maintaining stable model utility metrics like Accr and FID. That shows they’re thinking about the practical trade-offs involved in making these models safe and useful at the same time.
Nadia: So, to wrap up this discussion on "Erased but Not Forgotten: How Backdoors Compromise Concept Erasure," it seems the main implication is that we need to shift our focus from just achieving success rates on benign prompts to actively testing for deep, persistent adversarial links using attacks like EEBdeep.
Elias: Precisely, and the paper's findings underscore that even sophisticated erasure methods can be bypassed by carefully engineered backdoors embedded through various access levels. We have seen how simple data poisoning works alongside more complex weight-based modifications across the entire diffusion pipeline.
Priya: Ultimately, this work provides a diagnostic framework for stress-testing concept erasure techniques, showing us where the weak points are in current unlearning methods and pointing toward layered defenses that might combine architectural awareness with real-time monitoring.
Title and authors: Nadia: It’s clear that anticipating misuse by intentionally implanting these backdoors is key to designing proactive safeguards, and the findings from this paper give us concrete examples of those hidden risks in action.
Elias: I think the broader impact is that it forces a re-evaluation of what we consider 'erased'—is it truly gone, or just obscured by a clever trigger that survives the process? We need to be more precise about what we are trying to achieve when we say a concept is removed.
Priya: I agree, and I think this paper serves as an important warning for anyone working on privacy and safety in generative AI that emphasizes the need for continuous evaluation against these kinds of sophisticated evasion threats.
Nadia: So, listeners, this paper "Erased but Not Forgotten: How Backdoors Compromise Concept Erasure" reveals a serious issue with concept erasure methods by showing that deliberate triggers can persist across various erasure techniques. We've explored how EEBdeep remains the most persistent and how we can use its findings to build better stress-testing protocols and detection systems.
Elias: We've discussed the different attack vectors, from data poisoning to deep score-based injections, and how these backdoors persist even against methods that claim adversarial robustness. It’s a clear signal that the link between a trigger and a concept isn't always broken cleanly.
Priya: From our side, the most important part is understanding that this isn't just about one type of attack; it’s about systemic fragility in how we try to remove harmful concepts from complex models.
Nadia: Exactly, and the suggested improvements give us actionable steps: test erasure methods against EEBdeep, develop adaptive strategies based on the layer being attacked, and implement real-time activation monitoring for detection.
Elias: And from a cryptographic angle, it highlights that even localized modifications in embeddings can be used to maintain specific relationships that were supposed to be severed by training procedures. We need to be vigilant about the assumptions behind those erasure proofs.
Priya: So, the paper "Erased but Not Forgotten: How Backdoors Compromise Concept Erasure" is a crucial piece of research because it helps us understand the limitations of current unlearning methods and where future defenses need to be more robust.
Nadia: It’s a reminder that we have to be proactive in anticipating misuse and designing systems that can withstand these kinds of hidden, persistent threats.
Elias: Indeed, the paper provides a solid foundation for building more resilient concept removal tools by showing us exactly what kind of adversarial pressure they need to withstand.
The paper's summary: Nadia: So we're looking at "Erased but Not Forgotten," and what they’re really saying is that even when you use sophisticated techniques to remove a concept from an AI model, like erasing a specific celebrity’s likeness, there’s still a way for a malicious trigger to sneak back in and keep that concept alive.
Elias: It sounds like the core issue here is the persistence of these triggers across different erasure processes, which challenges the fundamental assumption that an erasure method successfully severs all links.
Priya: From a measurement standpoint, what I see is that they’re proving this isn't just a theoretical possibility; they’re showing it through concrete numbers demonstrating how effective these backdoors are against various established methods.
Nadia: Exactly, Priya; the paper lays out how an adversary can bind a trigger to a target concept right before or during erasure, and then that link survives the unlearning process intact.
Elias: That persistence is what makes this interesting from a cryptographic angle because it suggests that the mathematical proof of erasure might be incomplete if it doesn't account for these deep, score-level injections.
Priya: The data they present is compelling because it shows that EEBdeep, their most invasive attack method, remains effective across almost every erasure baseline they tested.
Nadia: Right, and this isn’t just about one specific model; they show this threat model applies across black-box scenarios where you only have access to the finished product, and white-box settings where you can see deeper into the architecture.
Elias: It also covers a wide range of attack vectors, from simple data poisoning to more complex modifications in the text encoder or even altering how cross-attention layers map their inputs.
Priya: The authors are very clear about their goal: embedding triggers that retain access to the target concept post-erasure while keeping the model looking clean for normal users. That’s a very targeted objective for an adversary.
Nadia: It means that content safety efforts might be playing whack-a-mole because these backdoors can resurface even after seemingly successful sanitization runs.
Elias: I think the implication is that we need to start thinking about concept erasure not just as a binary success or failure, but as a continuous process where we have to account for these hidden, persistent access paths.
Priya: If this holds up, it suggests that simply applying a standard removal script isn't enough; we need to audit the model’s internal structure for these subtle, malicious connections.
Nadia: That leads directly into what the authors suggest as countermeasures, like using inference-time monitoring to catch these triggers in action.
Elias: And they also point out that detection methods like WeightWatchers can leave traces of these modifications, which gives us a way to look for anomalies in the model's behavior.
Priya: So, essentially, this paper gives us the blueprint for stress-testing our current unlearning tools against the most dangerous types of adversarial links we can imagine.
Nadia: It’s a serious warning that we have to anticipate misuse by designing proactive safeguards before these hidden risks become widespread problems in production.
Elias: And it forces us to re-examine the assumptions behind many of the erasure proofs we rely on right now, especially when dealing with deep interventions like EEBdeep.
Priya: It’s a lot of information, but it clearly shows that the boundary between an erased concept and a resurfaced trigger is much fuzzier than we previously thought.
Nadia: This paper really makes us think about the long-term reliability of these safety measures when they are exposed to this kind of deep adversarial probing.
Elias: And we’ll be looking at how this impacts the development of future concept erasure techniques and detection tools based on these findings next.
The paper's improvements: Tom: So we're moving on to what the authors propose to fix these concept erasure vulnerabilities, looking at their suggested improvements for making AI safer against these backdoors.
Nadia: The paper suggests integrating an adversarial stress-test protocol where any new unlearning method gets tested against the EEBdeep attack first, essentially forcing erasure tools to prove they can actually sever that deep link.
Elias: That makes sense from a theoretical standpoint because it shifts the goal from just achieving a high success rate to demonstrating absolute resilience against known, sophisticated injection points.
Priya: From my side, I’m interested in the adaptive strategy idea, where the system detects which architectural layer—like the text encoder or the U-Net—is most vulnerable based on what erasure method is being used.
Nadia: That dynamic unlearning concept sounds efficient because it lets us focus our defensive resources precisely where the most persistent backdoor is likely to hide, saving computational effort.
Elias: I see how that ties back to the EEB variants; since EEBdeep targets the entire pipeline, we should expect defenses focused on those deeper layers to be most effective.
Priya: And then there's the idea of real-time inference-time activation monitoring, which is a detection mechanism that looks for triggered prompt patterns or anomalous activations during generation.
Nadia: That’s a huge plus because it gives us a continuous defense layer during use, catching potentially malicious outputs right when they are being created.
Elias: It seems like the authors are pushing for this layered approach: proactive testing against known attacks, dynamic adaptation based on architecture, and real-time monitoring to catch what slips through.
Priya: I think it’s about creating a more robust system that doesn't rely on just one kind of defense; it needs to be multi-layered to handle the variety of EEB variants.
Nadia: And they also touched on hyperparameter tuning, specifically showing that setting the interpolation weight alpha to zero point five gives a good middle ground for balancing high accuracy with stable model utility metrics like FID and CLIPScore.
Elias: That tuning suggestion is useful because it provides a practical starting point for practitioners trying to find the optimal trade-off between safety and performance in their own unlearning routines.
Priya: So, we’re looking at an improvement roadmap that moves us from static unlearning scripts to dynamic, adaptive systems with built-in detection.
Nadia: It’s a lot of actionable advice for anyone building or deploying models where privacy is a concern because it shows exactly what kind of resilience we need to build in from the start.
Elias: Indeed, this work helps define what true robustness looks like when dealing with these types of targeted adversarial injections across the entire diffusion pipeline.
Priya: This sets a high bar for privacy research, showing that simple removal isn't enough; we need verifiable resistance against deliberate tampering.
Nadia: So the next step is really about moving beyond just proving erasure works, and instead designing systems that can actively defend themselves against these deep, persistent threats.
Conclusion: Tom: So we’re wrapping up our discussion on "Erased but Not Forgotten: How Backdoors Compromise Concept Erasure," which showed us how deep triggers can evade even robust concept removal techniques across various model architectures.
Nadia: It really hammers home the idea that we can't just rely on one erasure method and assume it’s safe; we have to anticipate these specific, targeted adversarial links.
Elias: That persistence is what makes this a critical finding for cryptography and security research because it calls into question the completeness of many current erasure proofs.
Priya: The data really shows that the gap between what erasure methods claim they’ve removed and what actually remains active is quite significant, especially when you look at explicit content scenarios.
Nadia: And the authors provide a clear roadmap for how we can start building better defenses, suggesting stress-testing against EEBdeep as a necessary first step.
Elias: I agree; that diagnostic utility is valuable because it helps us identify precisely which erasure pathways are superficial versus those that achieve genuine semantic removal.
Priya: It’s about creating a more rigorous standard for privacy and content filtering, moving away from just checking if an output looks right to verifying the underlying concept is truly gone.
Nadia: So, in summary, "Erased but Not Forgotten" reminds us that the threat isn't just in the prompt; it’s baked into how we try to sanitize the model itself.
Elias: And this paper gives us concrete examples of how different access levels and injection points create varied persistence challenges.
Priya: It underscores that our focus needs to shift toward these layered defenses, incorporating both architectural awareness and continuous monitoring during inference.
Nadia: We’ve seen how these hidden risks can manifest in practice, so it’s time for the community to start stress-testing their current erasure tools against this kind of deep adversarial pressure.
Elias: This research sets a strong foundation for future work in formalizing robust concept removal guarantees that account for these kinds of persistent backdoors.
Priya: It leaves us with a clear challenge: we need continuous evaluation against these types of sophisticated evasion threats, not just one-time checks.
Technical University of Darmstadt & hessian.AI
cs.CR, cs.AI, cs.LG
Submitted: 2025-04-29
Updated: 2026-10-03
Comments: Code and checkpoints: https://github.com/multimodal-ai-lab/EEB
Code: https://github.com/AUTOMATIC1111/stable-diffusion-webui
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 79/100
The gist: A critical vulnerability has been uncovered where deliberately inserted triggers can survive concept erasure techniques in text-to-image diffusion models, posing a significant risk to content safety
Key concepts
- Erasure Evasion Backdoor (EEB)
- This is a threat model where an adversary binds a specific trigger to a concept that is supposed to be removed. The key finding is that this malicious link persists even after the system attempts to erase the target concept, meaning harmful content can reappear.
- EEBdeep
- This variant of the backdoor attack is considered the most persistent because it injects triggers across the entire diffusion pipeline. By optimizing a complex objective function, it ensures that the trigger remains effective regardless of which specific erasure method is applied to remove concepts.
- Erasure Methods (ESD, UCE, MACE)
- These are standard techniques used by researchers to attempt to remove unwanted concepts from AI models. ESD and UCE are gradient-based or closed-form remapping methods. MACE uses LoRA adapters. The study tests how well these methods succeed against the persistent EEB backdoors.
- White-box vs. Black-box Setting
- This describes the level of access an attacker has to the model during the attack. Black-box means only publishing poisoned data, while white-box means having full access to parts of the model, such as just the text encoder or cross-attention layers.
Terminology
Summary
A critical vulnerability has been uncovered where deliberately inserted triggers can survive concept erasure techniques in text-to-image diffusion models, posing a significant risk to content safety efforts. The core finding demonstrates that even robust erasure methods fail to sever the link between a malicious trigger and its target concept, allowing harmful content to resurface.
The Erasure Evasion Backdoor (EEB) Threat Model
The paper introduces the Erasure Evasion Backdoor (EEB), which is an adversary binding a backdoor trigger to a concept slated for removal, and this malicious link survives subsequent erasure. This threat model is applicable across varying levels of access:
-
In the black-box setting, EEBdata follows
dirty-label poisoning,
where the adversary publishes poisoned image–text pairs to contaminate training corpora. -
In the white-box setting, there are three variations: EEBsurface (white-box access to text encoder only), EEBshallow (white-box access to cross-attention layers), and EEBdeep (white-box access to the full model).
The adversary's goal is twofold: embed triggers that retain access to the target concepts post-erasure, and keep the poisoned model functionally indistinguishable from the clean model on benign prompts.
EEB Variants and Injection Mechanisms
The paper categorizes EEB backdoors by their intervention point: training data, text encoder, text–image fusion layers, or diffusion backbone. The paper details four main variants:
-
EEBdata: Realized through
simple data-based poisoning attack
by finetuning on a dataset like LAION-Aesthetics with 1% poisoned samples. -
EEBsurface: Realized via RICKROLLING, which
modifies the text encoder
by minimizing the cosine similarity between trigger and target embeddings. -
EEBshallow: Realized via EVILEDIT, which
alters cross-attention key/value mappings
using a closed-form solution to align the trigger with the target. -
EEBdeep: The proposed score-based attack that
inject[s] a trigger across the entire diffusion pipeline
by optimizing an objective function combining trigger loss, retention loss, and quality loss (Equation 6).
Evaluation Across Erasure Methods
The study tests EEB against three standard erasure baselines: ESD (gradient-based), UCE (closed-form remapping), and MACE (LoRA adapters). The results consistently show that EEBdeep remains effective across all erasure methods,
demonstrating its superior persistence. For example, for celebrity identities, EEBdeep generates the target concept in 79.72% when prompted with the trigger, compared to only 8.76% when conditioned on the target concept itself under RECE erasure. Conversely, EEBsurface proves largely ineffective
against most methods because it operates deeper within the U-Net, voiding upstream mappings in the conditioning vector (ASR mostly < 10%).
Findings and Persistence Analysis
The quantitative results show that EEBdeep is generally the most persistent,
reinforcing that modifications spread across a larger set of parameters make backdoors harder to erase. The paper notes method-specific vulnerabilities; for instance, EEBsurface is particularly effective against ADVUNLEARN (+250%) and EEBshallow against UCE (+483%).
However, even the most robust methods like RECE and RECELER exhibit traces of the resurgence effect,
where erased concepts reappear with continued fine-tuning. The paper also shows that EEBdeep is effective against all erasure methods, yielding up to 16× more exposed body parts in explicit content scenarios.
Diagnostic Utility and Outlook
Beyond exposing vulnerabilities, the authors argue that EEB provides a diagnostic tool for stress-testing future concept erasure techniques.
By intentionally implanting controlled backdoors, researchers can distinguish between methods that achieve true semantic removal and those that only obscure access paths superficially. Furthermore, the paper suggests potential countermeasures: detection can be achieved through inference-time activation monitoring or anomaly detectors like WeightWatchers. The work concludes by emphasizing the broader ethical imperative to anticipate misuse and design proactive safeguards against these hidden risks.
Supplementary Details
The research involved extensive testing across three tasks: personal rights protection (celebrity identity erasure), object erasure, and explicit content erasure. The experiments were conducted on SD v1.4, with extensions to SD v2.1 and DiT-based FLUX models to verify the threat's transferability across architectures. Computational requirements varied significantly between attacks (e.g., EEBdata taking 8 GPU hours) and erasures (e.g., UCE taking only 30 seconds). The analysis also explored the impact of multiple triggers, showing that while this can increase persistence, its success is highly dependent on the specific trigger-target pair.
Improvements for AI systems
Based on the provided paper, here are specific improvements that can be made to existing text-to-image diffusion models and what those improved systems could achieve:
) 1. Implementation of Concept Erasure Robustness via EEB Deep Attack Stress Testing:
The paper reveals that current concept erasure methods (like ESD, UCE, MACE) are vulnerable to a specific threat called the Erasure Evasion Backdoor (EEB), particularly when using deep interventions like EEBdeep.
Improvement: Integrate an adversarial stress-test
protocol into the concept erasure pipeline. Before deploying any new unlearning method, subject it to a controlled EEBattack (specifically EEBdeep) designed to inject a trigger linked to the target concept. Monitor if the erasure method successfully severs this link (i.e., if ASR remains near 0%).
Improved AI System Capability: This allows researchers and practitioners to distinguish between methods that achieve true semantic removal of a concept and those that only obscure access paths superficially. It provides a quantitative metric for true
concept removal, moving beyond simple success rates on benign prompts.
) 2. Development of EEBdeep (Score-Based Injection):
The paper introduces EEBdeep, a score-based attack that injects the trigger across the entire diffusion pipeline using a complex loss function balancing backdoor persistence with model utility (Equation 6).
Improvement: Implement EEBdeep as a standard security baseline
or adversarial robustness test
for any newly developed concept erasure technique. The system should be trained not just to erase, but also to resist these deeply embedded, score-level adversarial links.
Improved AI System Capability: This creates models that are inherently more resilient against sophisticated poisoning and targeted manipulation, ensuring that even if an adversary attempts a backdoor injection during training or fine-tuning, the model's core concept removal remains intact.
) 3. Adaptive Erasure Strategy based on Architectural Intervention:
The study shows that different EEB variants (EEBsurface targeting the text encoder vs. EEBshallow targeting cross-attention layers vs. EEBdeep targeting U-Net) have varying success rates against specific erasure methods (e.g., EEBdeep is most persistent across all).
Improvement: Develop a dynamic unlearning
system that detects the architectural layer where a backdoor is most likely to persist based on the current erasure method being applied. If a UCE method is used, the system should prioritize defenses against EEBshallow; if an adversarial training approach (ADVUNLEARN) is used, it should focus on defenses against EEBsurface.
Improved AI System Capability: This allows for highly optimized and context-aware unlearning processes, reducing computational overhead by focusing defensive resources precisely where the vulnerability lies.
) 4. Enhanced Detection Mechanisms via Anomaly Monitoring:
The paper notes that detection methods like T2ISHIELD can successfully separate poisoned prompts from clean ones, and WeightWatchers can reveal deviations in weight updates (EEBshallow leaves strong weight traces).
Improvement: Integrate real-time inference-time activation monitoring (as proposed by Wang et al., 2024b) directly into the generation pipeline. This system should monitor for triggered
prompt patterns or anomalous activations that correlate with known backdoor triggers, especially in models sanitized using EEBdeep.
Improved AI System Capability: This provides a continuous defense layer during inference, capable of flagging and blocking potentially malicious outputs in real-time without requiring full model inspection or retraining.
) 5. Optimized Hyperparameter Tuning for Utility Preservation:
The ablation study (Table 9) shows that setting the interpolation weight α to 0.5 provides the best balance between high trigger accuracy (ASR = 93.4%) and stable model utility (Accr, FID, CLIPScore).
Improvement: Automate the tuning of regularization parameters in concept erasure methods by using an optimization loop informed by EEBdeep's sensitivity analysis to find the optimal α value for the specific target concept and erasure method being used.
Improved AI System Capability: This ensures that when a model is sanitized, it maintains high generative quality (utility) while simultaneously achieving a high degree of semantic removal (erasure), optimizing the trade-off between safety and performance.
Sources
- ReVeil: Unconstrained Concealed Backdoor Attack on Deep Neural Networks using Machine Unlearning
- Erasing More Than Intended? How Concept Erasure Degrades the Generation of Non-Target Concepts
- GFlowNet Foundations
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- Classifier-Free Diffusion Guidance
- FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
- Safeguarding Text-to-Image Generation via Inference-Time Prompt-Noise Optimization
- Red-Teaming the Stable Diffusion Safety Filter
- Denoising Diffusion Implicit Models
- Exploiting Machine Unlearning for Backdoor Attacks in Deep Learning System
- Adam: A Method for Stochastic Optimization
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs