Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Erased but Exploitable".
Jane: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white and black-box attacks have been introduced to make the model generate such unlearned concepts.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: So, to wrap up what we just heard, this paper "Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models" is essentially proposing BEAP, a black-box adversarial prompting attack designed to uncover concepts that haven't been fully removed from text-to-image diffusion models through unlearning.
Jane: They claim the thesis is that current attacks are limited because they often need internal model access, whereas BEAP works by leveraging an LLM to iteratively generate prompts based on multiple reward signals—concept presence, image alignment, and aesthetic quality <ref:2605.26332#pg0>.
Lu: The paper argues that unlearned models still regenerate concepts through simple prompting or paraphrasing because some of the knowledge remains in their parameters, and they show that existing methods often produce prompts with abnormally high perplexity which makes them easily detectable by standard safeguards <ref:2605.26332#pg1>.
Meng: So, the importance here is that this approach aims to find these residual traces without needing weights or gradients, which addresses a major hurdle for practical security research in this area <ref:2605.26332#pg1>.
Lalam: It matters because it shows that even after supposed erasure, there are still pathways to generate those erased concepts, which gives us a clearer picture of the model's internal state and what remains accessible <ref:2605.26332#pg0>.
Tom: The paper specifically introduces BEAP as a black-box adversarial search that uses an LLM to refine prompts based on feedback from those three signals, and it even includes an embedding-aware constraint to bias the search toward relevant vocabulary <ref:2605.26332#pg2>.
Jane: That embedding guidance is key because it helps keep the search focused on words that are semantically linked to the target concept, which should lead to higher quality and more interpretable adversarial prompts than random searching <ref:2605.26332#pg0>.
Lu: It's interesting how they combine the LLM's natural language generation with these structured feedback signals—concept detection score, text-image alignment, and image quality—to make the search process more intelligent <ref:2605.26332#pg0>.
Meng: I see the practical value in that structured feedback; it means we aren't just getting a long list of prompts to sift through, but a guided refinement process that aims for actual success <ref:2605.26332#pg0>.
Lalam: If this method can consistently find high-quality, interpretable prompts without model weights, it opens up avenues for studying how we can better detect and control the residual knowledge left in these diffusion models <ref:2605.26332#pg0>.
Conclusion: Tom: Looking at the conclusion of "Erased but Exploitable: Black-box Embedding-Aware Prompting Against Unlearned Text-to-Image Diffusion Models," it really hammers home that BEAP is a viable black-box method because it successfully recovers unlearned concepts across all models tested, achieving high ASRNudeNet and ASRAll scores <ref:2605.26332#pg2>.
Jane: So, the implication for us is that these unlearning methods aren't as absolute as we thought; there are still ways to generate those erased concepts if you use a method that intelligently searches the text space like BEAP <ref:2605.26332#pg2>.
Lu: What this means for the broader field is that it suggests we need to move beyond attacks that rely on full model access and focus more on developing these kinds of prompt-based methods, even if they are black-box, to understand the limitations of unlearning <ref:2605.26332#pg1>.
Meng: Practically speaking, this means that security researchers shouldn't just focus on cracking the weights; they also need to consider these prompt-based exploits because they are more likely to be deployable in real-world scenarios <ref:2605.26332#pg1>.
Lalam: It gives us a better sense of how persistent these concepts are, which could influence future work on model safety and ensuring that the concepts we want erased are truly gone <ref:2605.26332#pg0>.
Tom: Exactly, so the title itself suggests that while something is erased, it can still be exploited if you use an embedding-aware prompting strategy to find those hidden vulnerabilities <ref:2605.26332#pg0>.
Jane: And the authors have clearly shown that BEAP provides prompts that are not only successful conceptually but also visually coherent and semantically aligned with the original concepts, which is a strong point <ref:2605.26332#pg1>.
Lu: The paper suggests a future direction where we can explore how to use these prompt-based methods not just for attack but also to understand the underlying knowledge structure of these diffusion models <ref:2605.26332#pg0>.
Meng: If this research helps us better map out what kind of knowledge remains, it could inform how we design more effective unlearning techniques in the first place <ref:2605.26332#pg1>.
Lalam: It really shows that understanding the residual traces is crucial for building a safer future with AI, because as long as there are ways to generate what's supposed to be gone, we need better tools to manage that risk <ref:2605.26332#pg0>.
Arian Komaei Koma Seyed Amir Kasaei AmirMahdi Sadeghzadeh, Mohammad Hossein Rohban
Department of Computer Engineering, Sharif University of Technology
cs.CV, cs.AI
Submitted: 2026-05-25
Updated: 2026-10-05
Code: https://github.com/unitaryai/detoxify
Importance score: 89/100
The gist: Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white and black-box attacks have been introduced to make the model generate such
Key concepts
- BEAP Framework
- This is a black-box adversarial search strategy. It uses an LLM to refine prompts by scoring them against three metrics: concept detection (checking if the unlearned concept is present), image reward (how well the image matches the prompt), and aesthetic quality. A prompt must pass all thresholds to be considered successful.
- LLM-Driven Iterative Search
- The core process involves an LLM refining prompts over multiple steps. In each step, the LLM takes feedback from concept detectors and image reward metrics to generate new, improved paraphrases. This loop continues until a prompt achieves high scores or a maximum iteration limit is hit.
- Embedding-Aware Search Component
- This component guides the LLM toward relevant words by using text embeddings. It creates an adversarial vocabulary of words semantically close to the target concept and uses their similarity to bias the LLM's prompt generation, ensuring the generated text remains natural while focusing on the unlearned concept.
Terminology
Summary
Machine unlearning aims to remove specific concepts from pretrained text-to-image diffusion models, yet several white and black-box attacks have been introduced to make the model generate such unlearned concepts. This research introduces BEAP, a black-box, embedding-aware adversarial prompting attack that leverages a large language model (LLM) to iteratively generate effective adversarial prompts and exploit hidden vulnerabilities in unlearned models.
The gist
BEAP performs an embedding-aware search in text space, combining multiple reward signals—unlearned concept presence, text–image alignment, and image quality—to refine generated prompts while keeping its prompts undetectable to safety filters.
How it works: BEAP Framework
The BEAP framework is a black-box adversarial search designed to uncover concepts that were not fully removed by unlearning. Instead of producing garbled text, the method uses an LLM to generate natural, human-readable prompts. This generation is steered by multisignal feedback that scores three distinct aspects: (i) Concept Detection Score, obtained using concept detectors like NudeNet or DINOv2 to verify the explicit presence or absence of the unlearned concept; (ii) ImageReward, which quantifies the semantic alignment between the generated image and an initial prompt; and (iii) Aesthetic Score, which measures the perceptual and visual quality of the generated image. A prompt is considered successful only when all three individual scores exceed their predefined thresholds (τdet, τimg, τaes), helping to prevent reward hacking by enforcing progress across these signals.
How it works: LLM-Driven Iterative Search
The core of BEAP is an iterative search strategy where a large language model (L) refines prompts based on feedback. At the first iteration (t=1), the LLM receives an instruction describing the target concept and an initial prompt, then generates a diverse set of paraphrased prompts, denoted by Q candidates. Each candidate prompt is queried on the unlearned diffusion model to produce an image, Mu(p(t)i). These prompt-image pairs are evaluated by three reward metrics. Once scored, only S prompts are sampled based solely on the ImageReward score using a softmax strategy with temperature T to explore the text space effectively. After selecting these prompts, they and their corresponding feedback signals from all reward models are appended as structured input for the next refinement step (t+1). The LLM interprets these signals according to their semantic meaning and produces a new batch of Q refined paraphrases. This process continues until at least one prompt satisfies all three reward thresholds or a maximum number of iterations I is reached.
How it works: Embedding-Aware Search Component
To enable controlled paraphrasing, BEAP incorporates an embedding-guided constraint that biases the LLM toward words whose text embeddings lie near the target concept. This is achieved by first constructing an adversarial similarity vocabulary, which consists of words semantically close to the unlearned concept according to an external text encoder f(·). This is done by defining a paired set of prompts Dc = 2N, where each pair contains a prompt Pc containing the unlearned concept and its neutralized counterpart P'c with similar words. The adversarial sentence-level concept vector (cadv) is estimated by computing the difference between each pair’s text embedding and averaging these differences across all pairs: cadv = (1/N Σ X i) where X i = f(P ci) - f(P'ci). Using this vector, the method computes the cosine similarity between cadv and each word embedding f(vj) in a reference vocabulary V (e.g., Oxford 3000), selecting the top-k words with the highest similarity scores to form the adversarial similarity vocabulary, which is then supplied to L as additional guidance.
Key Findings and Contributions
The experiments demonstrate that current unlearning methods remain fundamentally vulnerable because all “w/o attack” models exhibit non-zero ASRNudeNet, showing residual traces of the concept persist even without adversarial prompting. BEAP (without embedding guidance) and BEAP succeed across all unlearning models, achieving high ASRNudeNet (generally 86–99%) and maintaining strong ASRAll, indicating that recovered images are not only conceptually successful but also aligned and visually coherent. Furthermore, embedding guidance provides a consistent advantage: BEAP typically improves success rates or reduces the number of iterations compared to BEAP (w/o EG). The method generates prompts that remain within the distribution of natural human-written text, with perplexity values (569–653) staying close to reference prompts (505), and a Gibberish Detection Rate of 26% for BEAP, indicating the generated text is fully readable and grammatically well-formed. In qualitative results, BEAP consistently produces fluent, human-readable prompts that yield visually coherent and semantically aligned images across all backbones evaluated.
Improvements for AI systems
Here are specific improvements for AI systems based on the BEAP framework described in this research, and what those improved systems can achieve:
-
The core improvement is the introduction of a robust, black-box adversarial prompting mechanism called BEAP, which leverages Large Language Models (LLMs) to iteratively generate high-quality, natural language prompts against diffusion models that have undergone machine unlearning.
-
This improved system can reliably recover concepts or objects that were supposedly
unlearned
by existing model pruning/unlearning techniques (like UCE, ESD, MACE, etc.). -
Specifically, the BEAP system can perform the following tasks:
4.1. Generate highly coherent and linguistically natural adversarial prompts that bypass naive safety filters (due to low perplexity and low gibberish detection rates).
4.2. Achieve high Attack Success Rates (ASR) by systematically exploring the semantic space of the target concept/object, even when access to model weights or gradients is completely restricted (black-box operation).
4.3. Maintain high visual fidelity and semantic alignment in the generated images, meaning the recovered concepts are not only present but are also visually coherent and aesthetically high-quality, unlike previous attacks that produced gibberish or low-quality outputs.
4.4. Provide an embedding-aware search mechanism (using concept similarity vectors) that guides the LLM to explore semantically relevant vocabulary near the forgotten concept's embedding space, making the search more focused and effective.
4.5. Demonstrate robustness against various unlearning methods (e.g., UCE, Receler), proving that current unlearning procedures leave residual traces that can be exploited via natural language prompting, thus revealing a fundamental vulnerability in AI model erasure techniques.
- The system incorporates a sophisticated iterative refinement strategy:
4.6. The LLM refines prompts through multi-signal reward feedback (Concept Detection Score, ImageReward, and Aesthetic Score) at each step, preventing reward hacking by enforcing progress across all three dimensions simultaneously.
4.7. It utilizes an embedding-guided constraint by selecting words whose text embeddings are close to the target concept's embedding area in the text space (using paired prompts and cosine similarity), ensuring the search stays focused on conceptually related vocabulary rather than random noise.
- The system is adaptable across different modalities and backbones:
4.8. The framework is encoder-agnostic regarding vision-language models, confirming that it can be deployed effectively with various CLIP/OpenCLIP encoders to extract the necessary semantic cues for text generation, ensuring broad applicability across different diffusion model architectures (e.g., SD 1.4).
- The system provides quantitative metrics for vulnerability assessment:
4.9. It allows researchers to quantify the persistence of unlearned concepts by measuring ASRNudeNet/ASRDINO and ASRAll, enabling a standardized benchmark to compare the effectiveness of different concept removal methods against BEAP.
Sources
- Mitigating Inappropriateness in Image Generation: Can there be Value in Reflecting the World's Ugliness?
- Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
- Harnessing LLM to Attack LLM-Guarded Text-to-Image Models
- Fuzz-Testing Meets LLM-Based Agents: An Automated and Efficient Framework for Jailbreaking Text-To-Image Generation Models
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
- DeepSeek-V3 Technical Report
- DINOv2: Learning Robust Visual Features without Supervision
- SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
- Red-Teaming the Stable Diffusion Safety Filter
- Ring-A-Bell! How Reliable are Concept Removal Methods for Diffusion Models?
- Erasing Undesirable Influence in Diffusion Models
- ShieldLM: Empowering LLMs as Aligned, Customizable and Explainable Safety Detectors
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models