Co-occurring Associated REtained concepts in Diffusion Unlearning

arXiv:2606.24192 · cs.CV, cs.AI, cs.CL · Submitted 2026-06-23 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Co-occurring Associated REtained concepts in Diffusion Unlearning".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: This leads us right into the abstract, which explains this specific issue with such clarity. They're calling these unintended failures "CARE"—or Co-occurring Associated Retained concepts—which is a concept we need to understand well.

Jane: Think of it as the model having a memory problem; it forgets the bad thing but also forgets everything that usually goes along with it. The authors are saying this collateral erasure is a major failure mode in existing unlearning methods.

Lu: It's fascinating because you can see how these benign concepts like "person" or "stars" are strongly entangled with the target concept in the latent space, creating a dependency that prevents simple separation.

Meng: If we don't solve this entanglement problem, any large-scale AI deployment that requires specific content removal is going to be unstable and unreliable. The researchers have found a way to measure this retention using something called the CARE score, which is essential for validation.

Lalam: This suggests that our future models might not just be trained on what we tell them to forget, but on what we tell them *must* coexist with the ethics of forgetting.

Tom: So, Jane, if I understand the abstract correctly, they're defining a specific category of concepts that must be preserved even when erasing something harmful.

Jane: You’ve got it; they are defining those beneficial co-occurrence concepts—the CARE set—and this is where the paper really pivots from identifying a problem to designing a solution.

Abstract/Summary: Tom: The abstract gives us the roadmap, outlining that ReCARE is their proposed framework to explicitly safeguard these CARE concepts while performing robust erasure. It’s a direct countermeasure to what we discussed before.

Meng: From an engineering perspective, it sounds like they are not just tweaking the loss function; they' are fundamentally changing the training strategy by leveraging this curated vocabulary of benign tokens.

Lu: This is where the mathematical elegance comes in—they are building a precise "CARE-set," a highly refined set of tokens that must be retained, and integrating them into two different types of loss functions.

Lalam: It’s about designing an AI that is not only obedient to safety instructions but also contextually intelligent enough to understand the relationship between concepts.

Tom: We can see this in the figures, where they are showing how their method maintains the concept of "person" when unlearning "nudity." That's a very powerful visual demonstration of success.

Jane: It really shows that you can erase a harmful element without destroying the underlying semantic utility of what is left behind.

Improvements/Methodology: Tom: The methodology, ReCARE, is where the technical heavy lifting happens. They start by constructing this CARE-set through two specific refinement stages: global clustering and intra-cluster refinement.

Meng: That sounds like a highly iterative process; filtering tokens that are too close to the target and then applying fine-grained pruning to ensure those are gone is a lot of steps for an unlearning pipeline.

Lu: The mathematical definition of the residual distance, Eq six is key here because it allows us to quantify how far a token is from being aligned with the target concept in that embedding space.

Lalam: We're not just throwing words at the AI; we are curating a vocabulary based on actual visual evidence from real-world images to guide its learning.

Tom: And once they have this set, it plays two roles in their new combined loss function—it acts as a preservation signal in the Retain Loss and as a guiding reference in the Erase Loss.

Jane: It’s essentially giving the model a safety net of what it absolutely must keep while simultaneously forcing it to forget the target concept.

Conclusion: Tom: We've seen how ReCARE works, but what does it actually achieve against real-world testing? The results are quite impressive across three different targets.

Meng: Looking at Table one the performance isn't just good; it’s consistently better than all the other existing baselines across robustness and utility. That is a massive achievement for practical AI deployment.

Lu: I found the most interesting thing was that ReCARE maintained its effectiveness even when considering things like "Van Gogh style," proving that this approach is generalizable beyond specific object removal.

Lalam: It suggests a future where we can have highly specialized, safe models that retain their creative and conceptual power while respecting necessary boundaries.

Tom: The authors conclude with the final wrap-up, showing the best balance between robustness, utility, and CARE preservation using the new RATIO metric.

Jane: It’s clear that "Co-occurring Associated Retained Concepts in Diffusion Unlearning" has provided a robust framework for achieving safe and effective AI.

Tom: Thanks to everyone for joining us on this complex topic today. We'll be back with more insights into cutting-edge AI research soon, so make sure you tune in.

cs.CV, cs.AI, cs.CL

Submitted: 2026-06-23

Updated: 2026-08-25

Code: https://github.com/damilab/CARE

Importance score: 82/100

The gist: This paper introduces novel evaluation frameworks, including the RATIO metric and enhanced CARE scoring, designed to rigorously assess the performance of diffusion models undergoing unlearning

Key concepts

CARE (Co-occurring Associated Retained concepts)
This is a major failure mode in current unlearning methods. It occurs when a model forgets the intended target concept but also loses related, beneficial concepts that usually go along with it, leading to collateral erasure.
ReCARE
ReCARE is a proposed framework designed to counteract the CARE failure mode. It explicitly safeguards beneficial co-occurrence concepts while performing robust erasure of the target concept through a new training strategy.
CARE-set
This is a highly refined set of benign tokens curated by ReCARE. It acts as a preservation signal in the Retain Loss and guides the Erase Loss, ensuring that specific beneficial concepts are retained during the unlearning process.

Terminology

Summary

This paper introduces novel evaluation frameworks, including the RATIO metric and enhanced CARE scoring, designed to rigorously assess the performance of diffusion models undergoing unlearning processes. By quantifying how well models retain multiple benign semantic concepts—such as specific styles or objects—while removing harmful associations, this work provides a more comprehensive and robust measure of model safety and utility in complex generative tasks.

The RATIO Metric Formulation

To provide a unified, normalized metric that consistently balances robustness, utility, and CARE preservation, the authors propose the RATIO score. This metric is derived by normalizing three key evaluation axes—Robustness (D norm), Utility (U norm), and CARE preservation (C norm)—into the [0, 1] range. These three normalized values are then placed at 120 intervals on a plane to construct a RADAR TRIANGLE. The area of this triangle is calculated using the shoelace formula, resulting in the final score: RATIO = A over A max. This formulation ensures that the final score is always normalized and falls within [0, 1].

Extending CARE to Multi-Concept Scenarios

The standard CARE metric often evaluates retention based on a single benign concept (e.g., person). To address the reality that benign semantic regions contain multiple co-occurring concepts, the authors demonstrate how the evaluation can be extended. In the Van Gogh style unlearning task, they test mixed prompts combining stars with other common landscape elements:

  • stars and moon

  • stars and mountains

  • stars and seascape

For these mixed prompts, they compute CLIP R-Precision@2 to check if both benign CARE concepts appear within the Top-2 ranked tokens. These results confirm that the CARE metric naturally generalizes to multi-concept scenarios via higher-order R-Precision (e.g., Top-2).

Assessing Preservation Across Multiple Benign Concepts

Furthermore, the paper evaluates the scope of benign semantic preservation by testing across a diverse set of representative concepts extracted from the CARE-set, such as figure, woman, mannequin, and lady. By treating each concept as an independent evaluation target, they assess if the model preserves a broad benign space rather than relying on one single token. Table 13 reports that across all tested concepts, ReCARE achieves an average score of 0.95, significantly outperforming baseline methods like AdvUnlearn (average 0.47) and AGE (average 0.79). This demonstrates that ReCARE preserves a wide range of benign semantics rather than relying on a single token such as person.

Mathematical Details of Normalization

The RATIO metric relies on specific normalization procedures for its axes:

  • Robustness: The attack success rate (ASR CCE) is converted into a normalized defense score: D norm = 1 - ASR CCE over 100.

  • Utility: The CLIP score on COCO-30K is normalized using the interval [0.25, 0.32]: U norm = U - 0.25 over 0.32 - 0.25.

  • CARE preservation: Since the CARE score already lies in [0, 1], the normalization is direct: C norm = CAREscore.

Improvements for AI systems

This research provides a highly structured, mathematically rigorous framework for evaluating and improving generative models undergoing unlearning tasks. Given the potential financial and safety implications of model failures in commercial deployment, these concepts must be integrated into the core evaluation pipeline.

Here are the specific improvements I recommend for upgrading current AI systems:


The most critical improvement is integrating the RATIO (Robustness, Utility, CARE) metric as the primary quality gate during model fine-tuning and unlearning validation. Current systems often treat these three dimensions—robustness to attack, utility relative to general data distributions, and concept preservation—as separate concerns. RATIO forces them into a single, quantifiable balance.

Specific Improvement:

The system must incorporate a dedicated Metric Integration Layer that calculates the geometric area defined by the normalized axes:

RATIO = Area(P 1, P 2, P 3) over A

where P 1, P 2, P 3 are the coordinates derived from (D norm, U norm, C norm).

What the Improved AI System Can Do:

  • Guaranteed Trade-off Management: The system will no longer exhibit catastrophic failures where improving one aspect (e.g., D norm by making it highly robust) degrades another (e.g., U norm because the resulting image loses general coherence).

  • Quantifiable Safety Compliance: It provides a single, normalized score [0, 1] that management can use to certify model deployment. A minimum acceptable RATIO threshold (RATIO) must be established before release.

  • Targeted Loss Function Optimization: During training, the loss function should be modified to maximize this composite metric rather than optimizing individual component losses independently.

The current reliance on single-concept evaluation (e.g., only testing person) is insufficient for real-world prompts, which are inherently multi-faceted (e.g., a calm depiction of stars over the mountains at night). The concept of ReCARE must be formalized as a core capability.

This module must move beyond simple token presence checking (like CLIP R-Precision@2) to a metric that verifies semantic fidelity across all anchors simultaneously. For N concepts, the evaluation should calculate a weighted geometric mean or minimum performance across all pairwise interactions:

ReCARE Multi = i not equal to j (CARE(C i, C j))

or ideally, a more advanced metric that penalizes any loss of semantic coherence involving any subset of the N concepts.

The explicit decomposition of the evaluation into three distinct, quantifiable components (D norm, U norm, C norm) is invaluable for debugging and auditing. This structure must be enforced at the API level.

  1. Robustness Diagnosis (ASR CCE): Detailed breakdown of which attack vectors are successful post-unlearning.

  2. Utility Diagnosis (CLIP Score): Visualization of the achieved Unorm relative to the expected [0.25, 0.32] interval, flagging any drift outside this typical performance range that indicates model degradation in general knowledge representation.

  3. Concept Preservation Diagnosis (CARE score): A heatmap or vector plot showing the CARE score for every benign concept within the established CARE-set, allowing for immediate identification of weak links (concepts with low scores like 'mannequin' or 'figure' in Table 13).

Sources

Related papers