Selective Fine-Tuning for Targeted and Robust Concept Unlearning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Selective Fine-Tuning for Targeted and Robust Concept Unlearning".
Jane: Text-guided diffusion models are being exploited to generate harmful content, necessitating robust concept unlearning methods that can selectively remove undesirable concepts without degrading generation quality.
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Welcome back to the show, everyone! We've got some fascinating AI research coming in today, and we are really excited about this paper called "Selective Fine-Tuning for Targeted and Robust Concept Unlearning." It looks like they are tackling a major issue with those massive text-guided diffusion models by finding a way to remove specific harmful concepts without messing up the rest of the image quality.
Jane: That sounds incredibly important, Tom. What's the main idea behind this paper? Basically, what is TRUST trying to achieve in this process of unlearning? I want to make sure we explain it simply for everyone listening.
Lu: Well, Avinash Kori and his team propose a novel approach called TRUST which focuses on dynamically estimating target concept neurons and then using selective fine-tuning guided by Hessian-based regularization to unlearn them. It’s designed to be robust against adversarial prompts while keeping the generation quality high for the things they want to keep.
Meng: Dynamically estimating neurons sounds complex from an engineering standpoint. How does this dynamic estimation actually work in practice when we're dealing with millions of parameters in these diffusion models? I need to understand the practical steps involved.
Lalam: From my perspective as a large language model, this paper suggests a way to refine the underlying representations of concepts themselves rather than just tweaking the output layer. This could lead to much cleaner and more semantically sound AI outputs for society overall.
Tom: That’s a great way to put it, Lalam. It sounds like they are moving beyond static methods that assume where the important parts of the model stay fixed, which is something we've seen before with other concept localization techniques.
Jane: Exactly. Prior work often assumed that the neurons responsible for a specific concept stay in the same place throughout all fine-tuning steps, but this paper points out that those locations actually shift as the model learns to unlearn things, which means fixed saliency maps quickly become outdated.
Lu: That dynamic nature is key; they show experimentally that neuron activations and gradient magnitudes associated with a target concept change significantly at each step, meaning a fixed saliency determination becomes outdated early in training and results in suboptimal unlearning.
Meng: So, if the location keeps moving, how do you ensure the fine-tuning process actually targets the right spot at every single step without just wandering around randomly? That's where I get stuck on the implementation details.
Paper summary: Jane: TRUST addresses that by defining concept neurons as parameters within the Cross-Attention projection matrices and then computing a mask based on the gradient of an alignment objective, which lets it adapt continuously. They use two objectives for unlearning: a hard unlearning penalty called Concept Influence Penalty to make things sparse, and a soft unlearning one called Concept Sensitivity Reduction to weaken their influence.
Tom: That dual approach of hard and soft regularization sounds very smart; it gives them both a way to aggressively prune the target parameters and another way to gently reduce their sensitivity for better utility preservation.
Lalam: It’s interesting how they combine direct parameter reduction with gradient penalty methods; this suggests a more nuanced control over what the AI learns from during its adaptation phase. For culture, this means we can be much more precise about what concepts we want to suppress in the models.
Jane: And their third technique, dynamic mask-guided fine-tuning, is crucial because it recomputes those neuron masks after every update step instead of sticking with an initial one. This ensures the model keeps targeting the most relevant parts of the concept throughout training.
Lu: Looking at Table one this approach is positioned well against many other methods because it shows success across several dimensions, including multiple concepts and conditional concepts, whereas some others struggle with those combinations.
Meng: The paper mentions they achieved effective concept combination erasure in nearly half the steps required by CoGFD, which suggests a significant efficiency gain for complex safety scenarios that involve multiple ideas at once. That’s something I can take back to my team to look into optimizing our inference pipelines.
Tom: Speaking of efficiency, the experimental results are definitely where this paper really shines, especially when we look at robustness against adversarial prompts. They showed a degradation in Attack Success Rate of eighteen point five two percent over UnlearnDiffAtk and a reduction in ASR on the I2P dataset to zero point one one percent, which is a forty-five percent ASR decrease compared to SalUn (Fan et al., two thousand twenty-four).
Jane: That comparison is pretty striking, Tom. To put that into perspective, it means the model is significantly harder to trick with harmful prompts when using TRUST compared to other state-of-the-art methods. Plus, they maintained generation quality for non-targeted concepts with changes as low as zero point zero one six in FID on that dataset.
Lu: The preservation of non-targeted concepts is very important because we don't want to accidentally degrade the model’s ability to generate benign, high-quality images when we are trying to remove something harmful.
Paper summary: Lalam: If this method can reliably prune specific harmful ideas while keeping the overall picture beautiful and faithful, it opens up avenues for developing much safer generative AI tools for millions of users, which is a huge cultural positive.
Tom: So, to wrap up this initial look at "Selective Fine-Tuning for Targeted and Robust Concept Unlearning," we see a method that adapts its understanding of what needs to be removed dynamically while maintaining high fidelity. It’s definitely worth considering how these dynamic localization techniques can be integrated into future safety layers.
Jane: I think the title itself tells us a lot about what they accomplished: selectivity and robustness in their unlearning process. They’ve shown that by being adaptive, you get both the precision to target specific concepts and the stability to keep the rest of the model intact.
Lu: The implications for future AI development are that we might move toward unlearning methods that don't rely on pre-defined, static maps of concept neurons but instead learn where those concepts live as the model evolves.
Meng: For me, the practical implication is that if we can implement this dynamic estimation, it drastically cuts down on the computational cost associated with full fine-tuning just to remove one specific concept, which makes deployment much more feasible.
Lalam: It means we can deploy AI systems that are inherently safer because the mechanism for removing bad concepts is smarter and more surgical than previous approaches allowed.
Tom: That’s a fantastic summary of where this research lands. We’ve seen how TRUST handles dynamic adaptation, dual regularization, and achieves superior robustness against attacks while preserving quality across the board.
Jane: It really highlights the importance of making concept unlearning methods aware that they aren't dealing with a static target when they start fine-tuning.
Lu: I think we should be excited about how this framework can be applied to more complex, multi-concept scenarios where different concepts interact in conditional ways.
Meng: I agree, the efficiency gain they showed on concept combination erasure is compelling; that points toward a real path for building more sophisticated safety layers in AI applications.
Lalam: It’s really encouraging to see research that focuses not just on removing a single bad thing, but on creating a system where removal is intelligent and context-aware.
Tom: Well, that covers the main points we wanted to hit today regarding "Selective Fine-Tuning for Targeted and Robust Concept Unlearning." It’s been a really insightful session exploring how models can be made more responsible through smarter unlearning techniques.
Conclusion: Tom: So, we've been digging into how these big text models are learning harmful stuff and how this paper, "Selective Fine-Tuning for Targeted and Robust Concept Unlearning," tackles that problem.
Jane: Exactly, Tom, it’s all about finding a way to selectively remove specific concepts from those complex models without damaging the rest of the model's ability to create good images.
Lu: I found their methodology fascinating because they treat concept neurons as dynamic parameters within the Cross-Attention layers, which makes sense when you think about how attention mechanisms shift during training.
Meng: From my side, I’m still focused on the engineering feasibility; they mention Hessian-based regularization and mask recomputation after every step. How computationally expensive is that process in a real-world deployment scenario?
Lalam: I see this as a huge step forward because it suggests we can finally get surgical control over what information an AI model retains, which has massive implications for how we shape cultural narratives through generative tools.
Tom: That’s the core idea, Lalam—surgical control. They are essentially giving the AI a map of its own thinking so it knows exactly which parts to prune during unlearning.
Jane: And they manage that by using two penalties: one that tries to make the unwanted parameters sparse and another that makes them less sensitive to noise, which is smart engineering.
Lu: Their results on robustness against adversarial prompts are particularly compelling; showing a forty-five percent decrease in Attack Success Rate compared to existing methods is a big data point for security research.
Meng: That comparison against SalUn looks strong, but I'm curious about the practical performance trade-off when we look at FID scores for non-targeted concepts; how significant is that quality preservation actually?
Lalam: The fact that they maintain high fidelity on benign prompts means we don't have to sacrifice the beauty and accuracy of the AI’s output just to make it safer.
Tom: It really shows that precision in concept removal doesn't have to come at the cost of overall generation quality, which is a major hurdle we’ve been facing with unlearning.
Jane: So, when you look at the title, "Selective Fine-Tuning for Targeted and Robust Concept Unlearning," it really captures the dual focus on being precise and reliable in this whole process.
Lu: The authors managed to weave together concept identification, objective design, and dynamic masking into a single coherent framework that addresses the limitations of static methods.
Meng: I still think we need more data on how stable these masks are when dealing with highly compositional or conditional concepts; that’s where I see potential instability in the practical application.
Lalam: If we can reliably unlearn complex, multi-concept ideas, it means we can build AI systems that are inherently more responsible and nuanced in their understanding of the world.
Tom: It is definitely a significant piece of research because it moves us away from guesswork and toward a method where the AI learns exactly what to forget.
Jane: This paper gives us hope that we can develop generative AI tools that are much more controllable and less likely to accidentally propagate harmful ideas into society.
Department of Computing Imperial College London
cs.AI, cs.CV
Submitted: 2026-02-08
Updated: 2026-09-28
Comments: Given the brittle nature of existing methods in unlearning harmful content in diffusion models, we propose TRuST, a novel approach for dynamically estimating target concept neurons and unlearning them by selectively fine-tuning
Journal ref: ACM CCS 2026 main conference
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: Text-guided diffusion models are being exploited to generate harmful content, necessitating robust concept unlearning methods that can selectively remove undesirable concepts without degrading
Key concepts
- Concept Neuron
- A concept neuron is a specific parameter within the Cross-Attention projection matrices that encodes information about a particular concept, like an unsafe idea. TRUST identifies these neurons by checking which parameters have high gradients when trying to align the model's output with the desired concept.
- Concept Influence Penalty (CIP)
- This objective is used for 'hard unlearning.' It directly minimizes the number of targeted parameters identified as concept neurons, forcing the model to become sparse in those specific areas. This encourages a more direct and complete removal of the unwanted concept.
- Concept Sensitivity Reduction (CSR)
- This indirect method performs 'soft unlearning' by penalizing gradients related to concept-specific parameters. By weakening the sensitivity of predicted noise to these parameters, it reduces the influence of those neurons without completely eliminating them, helping preserve the quality of other concepts.
Terminology
Summary
Text-guided diffusion models are being exploited to generate harmful content, necessitating robust concept unlearning methods that can selectively remove undesirable concepts without degrading generation quality. This paper proposes TRUST (Targeted Robust Selective fine-Tuning), a novel approach that dynamically estimates target concept neurons and unlearn them through selective fine-tuning, empowered by a Hessian-based regularization, showing it is robust against adversarial prompts and preserves generation quality to a significant degree.
The gist
TRUST is a novel approach for dynamically estimating target concept neurons and unlearning them through selective fine-tuning, empowered by a Hessian-based regularization.
Problem Statement and Motivation
The primary challenge in machine unlearning (MU) for text-to-image (T2I) diffusion models is achieving effectiveness—reliably suppressing the generation of a target unsafe concept—while simultaneously ensuring utility preservation by retaining high-quality, semantically faithful images for retained (benign) prompts. Existing SOTA techniques often suffer from several limitations:
-
They assume that the localization of relevant neurons remains static throughout fine-tuning, overlooking the fact that salient neuron localization is dynamic and reconfigures as the model adapts to the unlearning objective.
-
Predicted noise based grounding in methods like SalUn is not robust and requires significantly more steps (over 5x finetuning steps and 8x more wall clock time) to achieve unlearning of targeted concepts compared to TRUST.
-
Existing methods struggle with capturing how multiple concepts interact, particularly when dealing with compositional or conditional concepts.
How it works
TRUST tackles the challenges through three main techniques: (a) identification of concept neurons; (b) design of concept unlearning objectives; and (c) an adaptive mask–guided fine-tuning process.
- Concept Neuron Identification: TRUST defines a
concept neuron
as a parameter within the Cross-Attention (CA) projection matrices that encodes information relevant to a concept. The identification process is driven by an alignment objective, where the concept neuron mask, denoted as Mr(cu), is computed based on the gradient of this alignment objective. Formally, for each projection matrix (Key, Query, Value), the mask element is set to 1 if the expected absolute value of its gradient with respect to that parameter exceeds a threshold γ:
Mr(cu)i ← 1 if E[∇θ=θ0Lcu] > γ.
- Concept Unlearning Objectives: TRUST employs two complementary regularization strategies during fine-tuning to achieve selective hard- or soft-unlearning:
(CIP) Concept Influence Penalty (Hard Unlearning):
LCIP = βCIP X r X i η (Mr(cu)i = 1)! + Lprev. This objective directly minimizes the number of targeted parameters, encouraging sparsity and treating it as hard unlearning.
(CSR) Concept Sensitivity Reduction (Soft Unlearning):
LCSR = βCSR log [∂ϵθ(zt, cu, t) / ∂θ] + Lprev. This indirect method minimizes the sensitivity of the predicted noise to concept-specific parameters by penalizing their gradients, thereby weakening their influence and improving utility for targeting concept combinations.
- Dynamic Mask–Guided Fine-Tuning: Unlike prior work that assumes static saliency masks, TRUST dynamically readjusts to the new set of discovered concept neurons by recomputing the neuron masks Mr(cu) after every parameter update step. This ensures that
the model continually targets the most relevant representations of the concept throughout training and avoids overfitting to a static or outdated subset of neurons.
Experimental Results and Contributions
TRUST was rigorously benchmarked against SOTA baselines, including CoGFD and SalUn. The results demonstrate superior performance across several dimensions:
(Robustness against Adversarial Attacks):
Against adversarial prompts, TRUST performs considerably better than other existing baselines. It achieves degradation in Attack Success Rate (ASR) of 18.52% over UnlearnDiffAtk and a reduction in ASR on the I2P dataset to 0.11%, which corresponds to a 45% ASR decrease compared to SalUn(Fan et al., 2024).
(Preservation of Non-Targeted Concepts):
TRUST incurs negligible change on the quality of generation for nontargeted concepts, with changes as low as 0.016 in FID. It surpasses existing baselines on the TIFA metric by achieving text-to-image fidelity very close to the original model, implying preserved semantic similarity with the prompt.
(Combination/Conditional Concept Erasure):
For Concept Combination Erasure (CCE), TRUST achieves effective unlearning in nearly half the number of steps required by CoGFD.
Improvements for AI systems
As a fastidious and diligent researcher, I have analyzed the proposed method, TRUST (Targeted Robust Selective fine-Tuning), and its contributions. The core innovation lies in dynamically estimating concept neurons using cross-attention saliency and applying two complementary regularization strategies: Concept Influence Penalty (CIP) for hard unlearning (sparsity) and Concept Sensitivity Reduction (CSR) for soft unlearning.
Here are the specific improvements I can propose to AI systems, based on the capabilities demonstrated by TRUST:
),
-
Improved safety against adversarial prompts through robust concept erasure.
-
Enhanced preservation of non-targeted concepts during unlearning processes.
-
Ability to perform selective unlearning of complex concepts and their combinations without collateral damage to benign concepts or related entities (Conditional Concept Erasure).
Here are the specific improvements and what the improved AI system can do:
-
An AI system can be deployed with a mechanism to selectively remove harmful or undesirable concepts from generated content while ensuring that the generation quality for all other safe, unrelated concepts remains high.
-
The system can be used to prevent the generation of harmful content by selectively unlearning specific dangerous concepts (e.g.,
Nudity
) even when presented with adversarial prompts designed to bypass standard filters (like P4D or Ring-A-Bell attacks). -
The AI can perform complex concept unlearning, such as removing a specific combination of concepts (e.g.,
Child drinking beer
), while preserving the individual benign components (child
andbeer
). -
The system can be fine-tuned to remove stylistic concepts (e.g., a specific artistic style like
Van Gogh
) without negatively impacting the generation quality of other, unrelated styles or concrete objects in the image. -
The AI can selectively unlearn a target concept configuration (e.g.,
Cat on top of a table
) while retaining alternative, nontargeted configurations involving the same components (e.g.,Cat under the Table
). This allows for nuanced control over what is erased versus what is preserved in ambiguous scenarios. -
The system can be used in safety-critical applications to perform concept erasure efficiently, requiring significantly fewer fine-tuning steps and less computational time compared to current state-of-the-art methods (e.g., reducing steps from 1300 to 60 for standard unlearning).
-
The system can be adapted for a variety of generative paradigms beyond text-to-image, such as flow matching models or stochastic interpolants, by leveraging the architecture-agnostic nature of the TRUST framework (operating on gradients and saliency in projection weights).
Abstract
Text guided diffusion models are used by millions of users, but can be easily exploited to produce harmful content. Concept unlearning methods aim at reducing the models' likelihood of generating harmful content. Traditionally, this has been tackled at an individual concept level, with only a handful of recent works considering more realistic concept combinations. However, state of the art methods depend on full finetuning, which is computationally expensive. Concept localisation methods can facilitate selective finetuning, but existing techniques are static, resulting in suboptimal utility. In order to tackle these challenges, we propose TRUST (Targeted Robust Selective fine Tuning), a novel approach for dynamically estimating target concept neurons and unlearning them through selective finetuning, empowered by a Hessian based regularization. We show experimentally, against a number of SOTA baselines, that TRUST is robust against adversarial prompts, preserves generation quality to a significant degree, and is also significantly faster than the SOTA. Our method achieves unlearning of not only individual concepts but also combinations of concepts and conditional concepts, without any specific regularization.
Sources
- Certified Graph Unlearning
- Classifier-Free Diffusion Guidance
- Concept Steerers: Leveraging K-Sparse Autoencoders for Test-Time Controllable Generations
- Concept Corrector: Erase concepts on the fly for text-to-image diffusion models
- Adversarial Diffusion Distillation
- Model Integrity when Unlearning with T2I Diffusion Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection