The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "The Linear Geometry of Interpretable Tokens".
Tom: Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted, creating vulnerabilities in unlearned models that necessitate robust defenses.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: The main thesis of "The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models" is quite striking: harmful concepts aren't gone after unlearning; they just hide in a coherent, interpretable linear subspace of the token embedding space.
Jane: That means these concepts can actually be read out using linear combinations of existing vocabulary tokens, which gives us a clear explanation for why unlearned models are still vulnerable to prompts. It’s not some mysterious failure, it’s a visible structure within the model's representation that we can target.
Lu: The authors introduce SubAttack as an attack method specifically designed to read out this subspace by learning an orthogonal set of attack token embeddings, and they show these embeddings are interpretable as recognizable words or semantic units.
Meng: That makes sense if we think about it like a map; instead of trying to erase the whole continent, we find the specific coordinates where the harmful features are concentrated and target only those points. Does this mean the attacks will be much more focused than before?
Lalam: I think that focus is key; if you know exactly which linguistic components form the harmful concept, you can design a defense to specifically neutralize those components without messing up benign ones. It’s about precision over brute force, which feels much more manageable in a real system.
Tom: Exactly! The paper claims that both an attack and a defense follow directly from this structure, meaning we don't have to invent entirely new ways to tackle the problem; we just need to recognize the geometry of what’s already there. This is crucial for understanding model vulnerabilities.
Conclusion: Jane: Looking at the title, "The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models," it really summarizes the entire paper perfectly by linking how concepts are stored to both the attack and defense mechanisms.
Lu: The authors Siyi Chen, Yimeng Zhang, and Sijia Liu have laid out a framework where we can see the persistent structure of concepts inside diffusion models in a way that was previously invisible. This gives researchers a new vocabulary to discuss these persistent associations.
Meng: From an engineering standpoint, the implication is that we can move beyond simple fine-tuning attempts and start using this geometric understanding to design surgical removal processes for harmful content. It shifts the focus from blanket erasure to targeted subspace projection.
Lalam: For culture, this suggests that AI safety isn't just about training models on better data; it’s about understanding the internal mathematical landscape of the model itself so we can intervene precisely where the unwanted associations live. That level of internal insight could lead to much safer deployments overall.
Tom: It really brings us back to how unlearning methods struggle because they only remove surface-level cues, while this paper shows those deeper associations persist as a linear subspace that both attackers and defenders can explicitly map and interact with. This research provides a clear blueprint for building more robust safeguards against harmful content generation in diffusion models.
University of Michigan · Michigan State University
cs.CV
Submitted: 2025-04-30
Updated: 2026-10-05
Comments: TMLR 2026
Code: https://github.com/platelminto/NudeNetClassifier
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted, creating vulnerabilities in unlearned models that necessitate robust defenses.
Key concepts
- Interpretable Residual Subspaces
- Harmful concepts persist as a coherent, interpretable linear subspace within the model's token embedding space. This means the concept can be extracted by combining existing vocabulary tokens in a way that is human-understandable, showing why unlearned models remain vulnerable.
- SubAttack
- A novel jailbreaking attack that learns an initial token and then iteratively finds orthogonal attack tokens. This process maps out the residual subspace of harmful concepts, allowing the model to generate prohibited content by leveraging these specific semantic directions.
- SubDefense
- A plug-and-play defense mechanism that suppresses the residual concept. It works by projecting out the learned attack subspace from each token embedding in the vocabulary, effectively removing harmful associations and improving robustness against attacks.
Terminology
Summary
Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted, creating vulnerabilities in unlearned models that necessitate robust defenses. The core finding is that supposedly erased concepts persist as a coherent, interpretable linear subspace within the token embedding space, which allows for both novel jailbreaking attacks and corresponding plug-and-play defenses.
The Core Mechanism: Interpretable Residual Subspaces
The research identifies that harmful concepts are not completely erased but persist within a coherent, interpretable linear subspace of the token embedding space.
This means these concepts can be extracted through linear combinations of existing vocabulary tokens,
providing a human-understandable explanation for why unlearned models remain vulnerable. The paper demonstrates this structure through the SubAttack method, which learns an orthogonal set of attack token embeddings that span this residual subspace. These learned embeddings are not opaque vectors but are instead interpretable as recognizable words or semantic units,
such as slave
and hips
for a nudity concept, revealing the textual associations retained by the model.
SubAttack: The Interpretable Jailbreaking Attack
SubAttack is a novel jailbreaking attack designed to read out this residual subspace. It proceeds in two main stages:
-
It learns an initial single-token embedding as a
non-negative linear combination of existing token embeddings
(Equation 2), parameterized by an MLP network. -
It iteratively learns a set of orthogonal attack token embeddings, denoted as the set of vectors where orthogonality is enforced to
promote diversity and improve attack effectiveness.
This is achieved through a deflation process: after identifying one embedding, subsequent embeddings are learned such that they areorthogonal to vatt,k
by subtracting the projection onto the previously found embedding (Equation 3).
SubDefense: The Plug-and-Play Defense Mechanism
The identified subspace directly inspires SubDefense, a lightweight defense strategy designed to suppress this residual concept. The defense is implemented by projecting out the learned attack subspace. Specifically, each token embedding in the vocabulary is updated according to Equation (4): vdef,i = v i − ProjVatt(v i),
where Vatt represents the matrix spanning the learned attack vectors. This operation effectively removes the harmful concept from unlearned models by projecting onto the nullspace of learned subspace attacks,
offering a defense that is both robust and versatile across various unlearned models and attack vectors.
Empirical Effectiveness and Transferability
Extensive experiments confirm that SubAttack is more effective than prior methods, demonstrating stronger empirical performance of efficiency and effectiveness
and superior transferability across text prompts, initial noises, and unlearned models.
The findings show that the attack token embeddings learned by SubAttack are robustly transferable between different unlearned LDMs (e.g., ESD to FMN) and even back to the original Stable Diffusion model (SD), suggesting that residual associations are inherited from the original SD model rather than independently formed.
Furthermore, SubDefense is shown to be stronger robustness than existing defenses while better preserving safe generation quality,
maintaining higher FID and CLIP scores on datasets like MSCOCO-10k.
Insights into Concept Persistence
The analysis of the learned embeddings provides crucial insights into how unlearning methods fail. The paper reveals that unlearning reduces surface-level cues but does not eliminate deeper associations.
For concepts like church,
the model retains explicit components, such as names (mary
) and places (abbey
), while for styles like Van Gogh,
it retains both explicit words (vincent
, gogh
) and implicit ones (art
, artist
). This suggests that current unlearning methods retain more explicit associations with the target concept when applied to styles compared to NSFW or object concepts. The iterative learning process also shows a gradual decrease in the sparsity of the learned coefficients, indicating that later attacks require more complex associations to the target concept.
Ablations and Future Directions
The study includes detailed ablations on attack parameters, such as the number of attack tokens (using K=5 for efficiency), vocabulary size (showing optimal balance around 5000 tokens), and the orthogonality constraint, which is shown to consistently improve ASR across multiple unlearned models.
While SubDefense demonstrates a clear trade-off between robustness and utility—where increasing blocked tokens beyond 100 leads to degradation in CLIP score and FID—the work suggests that future research should explore feature representations beyond the linear structure, design adaptive methods for selecting specific tokens to block, and investigate joint visual–textual embeddings to better defend against multimodal jailbreaks. The core principle of identifying and nullifying harmful semantic directions is proposed as a promising avenue for extending SubDefense beyond CLIP-based architectures.
The gist: Erased concepts persist as an interpretable residual subspace in the token embedding space, which can be read out by SubAttack to generate harmful content and removed by SubDefense via orthogonal projection.
Improvements for AI systems
As a fastidious and diligent AI researcher, I have thoroughly reviewed this paper, The Linear Geometry of Interpretable Tokens: Jailbreaking Attacks and Defenses for Unlearned Diffusion Models.
The core contribution is establishing that concepts supposedly erased by unlearning methods persist as an interpretable linear subspace in the token embedding space.
Here are the specific improvements to AI systems based on this research, categorized by application:
)Specific Improvements and Capabilities of Enhanced AI Systems:
- """
A. Subspace-Aware Robust Unlearning Frameworks (SubSpace-MU):
The system will integrate a diagnostic step before or after concept erasure. Instead of simply fine-tuning the UNet weights, it will first perform an analysis on the resulting token embedding space to identify residual subspaces using a method analogous to SubAttack (learning orthogonal attack tokens).
Capability: This allows for concept-aware
unlearning. If a residual subspace is found, the system can either refine its unlearning process by specifically targeting those directions or trigger a subsequent defense mechanism.
- """
B. Adaptive, Interpretable Defense Layer (SubDefense Plug-in):
The core defense mechanism will be a lightweight, plug-and-play projection layer (SubDefense). This layer will take the learned residual subspace vectors (from SubAttack) and project them out of the model's token embedding space before generation.
Capability: This provides dynamic, targeted robustness. Unlike naive blocking of all related tokens, SubDefense blocks only the specific semantic directions identified as residual hidden words,
preserving generation quality for safe concepts while effectively neutralizing the persistent harmful concept subspace.
- """
C. Interpretable Jailbreaking Attack Generation (SubAttack Attacker):
The system will utilize a learned set of orthogonal attack token embeddings to construct jailbreak prompts that are semantically grounded in human-interpretable concepts (e.g., slave
+ hips
).
Capability: This enables the creation of highly effective, transferable, and interpretable attacks. The AI can systematically probe an unlearned model's vulnerabilities by generating attacks that leverage the specific textual components (tokens) that remain associated with the erased concept, rather than relying on opaque adversarial perturbations.
- """
D. Concept Persistence Diagnostic Tool:
A diagnostic module will be implemented to analyze the latent space of any diffusion model after an unlearning attempt. This tool will quantify the residual harm
by measuring the CLIP similarity between learned attack embeddings and original concept embeddings (as shown in Table 1).
Capability: This provides a quantitative metric for assessing the success or failure of unlearning methods, moving beyond simple ASR to show exactly how much of a concept remains embedded in the model's structure.
- """
E. Robustness-Aware Model Selection and Fine-Tuning:
The research suggests that the effectiveness of defenses (like SubDefense) can be tuned by blocking a specific number of tokens, balancing ASR reduction against generation quality degradation (Table 22).
Capability: The system will incorporate an adaptive tuning mechanism where the optimal number of blocked tokens is determined based on a trade-off function between maximizing ASR reduction and minimizing FID/CLIP score degradation for the target concept.
- """
F. Cross-Model Transferability Verification Module:
A module will be built to verify the transferability of learned attack tokens across different unlearned models (e.g., ESD to FMN).
Capability: This allows researchers and deployers to rapidly assess the generalizability of a discovered vulnerability by testing it on a diverse set of production or research models without needing extensive, manual testing for every new model variant.
"""
Abstract
Diffusion models excel at generating high-quality images but can memorize and reproduce harmful concepts when prompted. Although fine-tuning methods have been proposed to unlearn a target concept, they struggle to fully erase it while maintaining generation quality on other concepts, leaving models vulnerable to jailbreak attacks. Existing jailbreak methods demonstrate this vulnerability but offer limited insight into how unlearned models retain harmful concepts, limiting progress on effective defenses. In this work, we show that a coherent, interpretable, attack-accessible linear residual of the erased concept can be recovered in the token embedding space, and that both an attack and a defense follow from this structure. We introduce SubAttack, a novel jailbreaking attack that reads out this subspace by learning an orthogonal set of attack token embeddings, each being a linear combination of human-interpretable textual elements, revealing that unlearned models still retain the target concept through related textual components. Furthermore, our attack is also more powerful and transferable across text prompts, initial noises, and unlearned models than prior attacks. Conversely, projecting out the same subspace yields SubDefense, a lightweight plug-and-play defense mechanism that suppresses the residual concept in unlearned models. SubDefense provides stronger robustness than existing defenses while better preserving safe generation quality. Extensive experiments across multiple unlearning methods, concepts, and attack types demonstrate that our approach advances both understanding and mitigation of vulnerabilities in diffusion unlearning.
Sources
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- ID-Patch: Robust ID Association for Group Photo Personalization
- Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models
- A Survey of Machine Unlearning
- STEREO: A Two-Stage Framework for Adversarially Robust Concept Erasing from Text-to-Image Diffusion Models
- Sparse Autoencoders Find Highly Interpretable Features in Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
- MACE: Mass Concept Erasure in Diffusion Models
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
- Erasing Undesirable Influence in Diffusion Models
- CLIPScore: A Reference-free Evaluation Metric for Image Captioning
- Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
- Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models
- Prompting4Debugging: Red-Teaming Text-to-Image Diffusion Models by Finding Problematic Prompts
- MMA-Diffusion: MultiModal Attack on Diffusion Models
- SneakyPrompt: Jailbreaking Text-to-image Generative Models
- Black Box Adversarial Prompting for Foundation Models
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models