Models Designed to Forget: Machine Unlearning via Key Deletion

arXiv:2603.15033 · cs.LG · Submitted 2026-03-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Models Designed to Forget".

Jane: Machine unlearning is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training samples.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re starting by looking at the title of "Models Designed to Forget: Machine Unlearning via Key Deletion," and it immediately tells us that these authors aren't just trying to patch up old unlearning methods; they are actually designing the model structure from the beginning to handle forgetting. The list of authors includes Laguna, Gonçalves, Vandenhirtz, Ryser, Cannistraci, and Vogt.

Jane: And what’s significant about this framing is how they propose moving away from post-hoc techniques that involve updating parameters during unlearning requests. They are suggesting that forgetting should be an inherent feature of the model architecture itself.

Lu: The authors support this idea by showing how separating the memory bank allows for a simple set-theoretic operation—just deleting an entry—to achieve zero-shot forgetting without needing any retraining or weight updates, which is quite a significant move.

Meng: It makes sense that they are shifting the responsibility of memorization away from the static weights and onto this externalized memory structure, especially when you consider deployment scenarios where data privacy needs change frequently.

Lalam: From what I’m seeing, the authors really focus on how this architecture allows for privacy-centric deployment by eliminating the necessity of storing raw samples or labels at the exact moment you need to forget something.

Tom: Right, and they back this up with recent findings that show visual memories can be used effectively even with billion-scale datasets, which gives them strong support for this entire approach.

Jane: It suggests that by using memory-augmented structures for the task of forgetting, they are creating a new category of models that are naturally compliant with changing privacy requirements.

Lu: I think the authors also point out that using external or decoupled memories has been explored before for things like boosting predictive performance or improving few-shot efficiency, so they are building on previous research in that area.

Meng: So, what they’re doing is taking those existing memory concepts and applying them specifically to the inverse task of data deletion, which is a clever way to use the technology.

Lalam: It really shows how foundational AI concepts can be repurposed to solve very specific constraints, like regulatory demands for data removal.

The paper's summary: Tom: So, summarizing "Models Designed to Forget: Machine Unlearning via Key Deletion," the authors propose a memory-augmented transformer where they decouple instance-specific memorization from the model weights by associating each piece of data with a learnable exemplar token.

Jane: Basically, they train the model so it learns to depend on this external bank of tokens for specific details, meaning when you want to forget something, you just remove that specific identifier from that bank.

Lu: The training involves injecting these tokens into the input sequence alongside image patches and optimizing both the main network parameters and these exemplar tokens at the same time using a cross-entropy loss.

Meng: The inference side is where things get interesting because for a new query, instead of looking at its original token, the model retrieves a set of nearest neighbors from this external memory bank based on key similarities.

Lalam: And then the final prediction isn't just one output; it’s an ensemble prediction calculated by weighting the logits from those retrieved tokens, which gives us a more stable final answer.

Tom: It really boils down to having the model look at its context and ask, "Which stored examples are most relevant right now?" which is a different way of making a decision compared to relying only on fixed internal weights.

Jane: This mechanism directly tackles the problem of knowledge entanglement in standard architectures by moving that memorization out into this external structure, which is what they call unlearning by design.

Lu: It suggests that the model doesn't need to keep every detail locked away in its weights; instead, it can use a flexible, accessible memory bank for instance-specific information when needed.

Meng: From an engineering standpoint, this means that as long as we have a reliable way to manage that external bank and the retrieval process works well, we can separate the model's core parameters from the data itself.

Lalam: I think this moves AI development toward systems where data is treated more like a dynamic resource instead of just static input for training.

The paper's improvements: Tom: Now let’s talk about how they improve upon existing ideas, and they introduce a stochastic pathway dropout to stop the model from becoming too reliant on just one way of processing the data.

Jane: And that dropout is combined with a Bernoulli sampling of a retrieval state, which decides whether the model uses the specific instance token or groups together neighboring tokens from memory. This adds another level of controlled variability to how it accesses that external memory.

Lu: The paper also introduces a Pathway Sensitivity Score, Ps, which helps them check the balance by testing performance when one pathway is replaced by learned null tokens,.

Meng: That Ps score is important for validation because it lets them see if the model actually keeps its utility while making sure both pathways—the raw image patches and the learned tokens—are contributing meaningfully.

Lalam: They also mention an ensembling strategy as being the most robust way to maintain this pathway balance while keeping that Average Gap low, suggesting ensembling is a critical structural element for controlling how well forgetting operates.

Tom: So when you put all those elements together, they are giving us a controlled mechanism to make sure that when we delete something, the model stays useful and accurate for what it was supposed to retain.

Jane: This control mechanism makes the process of removing data much more predictable than just letting the system wander around randomly during training.

Lu: The fact that they tested different key encoders like ViT, DINO-v2, and CLIP4 shows that this system is quite adaptable and doesn't need label information embedded in those keys to work correctly.

Meng: That adaptability is important because it means we aren't tied to one specific visual encoder for this unlearning technique; we can pick the best tool for the job.

Conclusion: Tom: So, wrapping up "Models Designed to Forget: Machine Unlearning via Key Deletion," the main implication is that we can now perform instance-specific data removal in a way that is both extremely fast and keeps privacy intact.

Jane: This capability means high-stakes environments, like clinical studies, can finally use AI with confidence knowing they have a surgical tool to surgically remove specific training examples without hurting the overall model performance on the data it keeps.

Lu: The structural flexibility shown across different key encoders opens up interesting avenues for how we can manage knowledge externally in even more complex AI systems.

Meng: For me, the practical impact is realizing that unlearning isn't just a theoretical exercise anymore; it’s a tool we can deploy to meet real operational needs in regulated industries.

Lalam: I think the overall direction of this work encourages us to treat data as a dynamic resource within AI systems rather than just static training input for future development.

Tom: Absolutely, so that’s where we are heading next, and I’m really looking forward to seeing how these design principles apply to other areas of AI research.

Jane: We definitely have a lot more interesting papers coming up, so stick with us as we keep exploring this fascinating space.

Department of Computer Science, ETH Zurich

cs.LG

Submitted: 2026-03-16

Updated: 2026-09-28

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 87/100

The gist: Machine unlearning is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training samples.

Key concepts

Unlearning by Design
Instead of removing data after training, this approach anticipates forgetting during the initial training process. It involves structuring the model so that instance-specific knowledge is stored externally rather than being permanently entangled in the main network weights, making removal straightforward.
MUNKEY (Machine UNlearning via KEY Deletion)
This is a method that stores each training example's unique information as a learnable token. By associating an image with a specific key and a learnable value, the model learns to rely on these external tokens for instance details, allowing specific data influence to be deleted by simply removing those tokens from memory.
Pathway Sensitivity Score (Ps)
This score measures how reliant the model is on different parts of its architecture—like image processing versus token processing. It ensures that when data is forgotten, the model doesn't collapse entirely by checking performance after replacing one pathway with null tokens, ensuring both utility and robustness.
Exemplar Memory Bank (M)
This is an external storage system where each training instance is mapped to a unique key-value pair. The keys are derived from the input data, and the values are learnable tokens. This bank acts as an organized repository for all specific instance information, separate from the core model weights.

Terminology

Summary

Machine unlearning is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training samples. The proposed method introduces unlearning by design, a novel paradigm in which models are directly trained to support forgetting as an inherent capability.

The gist: Unlearning corresponds to removing the instance-identifying key, enabling direct zero-shot forgetting without weight updates or access to the original samples or labels.

Problem Formalization and Paradigm Shift

Machine unlearning seeks to remove the influence of a specific subset of training data, denoted as the forget set, while preserving performance on the retain set. Current methods are largely post-hoc, relying on computationally intensive weight updates or retraining from scratch. The paper advocates for a paradigm shift towards unlearning by design, where forgetting is anticipated and explicitly enforced during training to streamline the operation. This approach addresses the fundamental structural flaw of knowledge entanglement in standard architectures, shifting the burden of memorization from static weights to an externalized memory bank.

MUNKEY: Machine UNlearning via KEY Deletion

The core contribution is MUNKEY, a method that decouples instance-specific data into an external exemplar memory bank M. This decoupling is achieved by associating each training instance with a learnable exemplar token, denoted as the key-value pair (ki, vi). The process involves:

  1. Defining an explicit exemplar memory bank M =

(ki, vi) N i=1 for all xi ∈ D.

  1. Extracting a fixed key ki = gϕ(xi) from a frozen encoder gϕ and defining a learnable exemplar token vi in R m.

  2. Injecting this into the input sequence to form zin, which is concatenated with image patch embeddings zi: zin = [zi, v′i].

Training with Stochastic Pathway Dropout

To prevent modality dominance—where the model over-relies on either raw image patches or exemplar tokens—MUNKEY employs a stochastic pathway dropout. A mask vector m = [γimg, γtok] follows a categorical distribution to sample whether to drop the image pathway or the token pathway. This is combined with a Bernoulli sampling of a retrieval state r, which determines whether to use the instance-specific token v′i or aggregate neighboring tokens vn from memory. The objective is optimized jointly:

E(xi,yi)∼D [l(fθ(zin), yi)], where θ and ψ are optimized alongside the exemplar tokens viN i=1.

Inference via Nearest Neighbors

At inference time, since the specific exemplar token vq for an unseen query xq is unavailable, the model retrieves K nearest neighbors NK(xq) from the updated memory bank Mu based on cosine similarity of kq = gϕ(xq) with stored keys. The final prediction yˆ is computed as a weighted average of logits: yˆ = P

j∈NK(xq) wj · fθ([zq, hψ(vj)]), where weights wj are determined by a softmax over key similarities.

Unlearning and Pathway Sensitivity

Unlearning itself is reduced to a simple set-theoretic operation: Mu = M (ki, vi) i ∈ Df. By removing the forget set entries from the memory bank, zero-shot forgetting is achieved because the backbone parameters θ and adapter ψ are trained to rely only on this externalized memory for instance-specific details. To ensure architectural balance and prevent pathway collapse, a Pathway Sensitivity Score (Ps) is introduced, quantifying information flow by evaluating model performance when one pathway is replaced by learned null tokens ∅. Models are selected based on minimizing Avg Gap while maintaining Ps above a threshold ξ = 0.3 to ensure robust forgetting and utility.

Empirical Validation

MUNKEY consistently outperforms all post-hoc baselines across natural image benchmarks (CIFAR-10, CIFAR-100, Tiny ImageNet) and medical datasets (BloodMNIST, PathMNIST, DermaMNIST). The results demonstrate that MUNKEY achieves the lowest Average Gap across random forget rates of 10% and 2%, confirming its superior balance between model utility and unlearning faithfulness. Furthermore, ablation studies show that while different aggregation strategies exist (e.g., Cross Attention vs. Ensembling), ensembling provides the most robust pathway balance with the smallest Avg Gap, suggesting it is a critical structural lever for unlearning. The method's runtime efficiency is negligible (≈ 0s) for unlearning, contrasting sharply with iterative post-hoc methods that require significant compute time.

Effect of Key Encoder and Neighborhood Size

Ablation experiments show that MUNKEY remains effective across different key encoders (standard ViT, DINO-v2, CLIP4), indicating it does not require label-informed keys to decouple instance-specific information.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed Rethinking Machine Unlearning: Models Designed to Forget via Key Deletion (MUNKEY). This paper introduces a paradigm shift from post-hoc unlearning (which requires expensive retraining) to Unlearning by Design using an externalized memory mechanism.

Here are the specific improvements and capabilities this system enables:


)

  1. The core improvement is the implementation of a modular architecture where instance-specific data is decoupled from static model weights via a learnable exemplar token stored in an external memory bank (MUNKEY).

  2. This system enables zero-shot unlearning via a simple set-theoretic operation: deleting the instance-identifying key from the external memory, without requiring any parameter updates or access to original samples/labels.

)

  1. The improved AI system can perform targeted data removal (e.g., GDPR right to be forgotten requests) in near-instantaneous time (≈ 0s unlearning time), bypassing the prohibitive computational cost of full model retraining or iterative optimization methods like gradient ascent or distillation.

)

  1. The system provides a superior balance between utility and forgetting through the use of a Pathway Sensitivity Score (Ps) and a rewind procedure for hyperparameter selection, ensuring that models are not over-reliant on either the raw image patches or the learned exemplar tokens.

)

  1. The improved AI system can be deployed in high-stakes environments (like medical imaging or clinical studies) where data privacy is paramount, as it maintains high predictive performance while guaranteeing surgical removal of specific training examples without degrading overall model utility on the retained data.

)

  1. The system offers enhanced interpretability by allowing researchers to pinpoint cornerstone samples (the influential neighbors retrieved for a decision), providing explicit evidence for predictions, which is critical for transparent case-based reasoning in healthcare or public policy applications.

)

  1. The architecture allows for the evaluation of different key encoders (e.g., DINO-v2, CLIP4) to determine that the system is robust and does not require label-informed keys, suggesting a more flexible and adaptable design space for future unlearning implementations across various vision tasks.

Abstract

Machine unlearning for vision models is rapidly becoming a practical requirement, driven by privacy regulations, data errors, and the need to remove harmful or corrupted training images. Despite this, most existing approximate unlearning methods tackle the problem from a post-hoc perspective. They attempt to erase the influence of targeted samples through parameter updates that typically require access to the full training data. This creates a mismatch with real deployment scenarios where unlearning requests can be anticipated, revealing a fundamental limitation of post-hoc approaches. We motivate unlearning by design, a novel paradigm for approximate methods in which models are directly trained to support forgetting as an inherent architectural capability. We instantiate this idea with Machine UNlearning via KEY deletion (MUNKEY), a memory-augmented transformer that decouples instance-specific memorization from model weights. Here, unlearning corresponds to removing the instance-identifying key, enabling zero-shot forgetting without weight updates or access to the original samples or labels. Across natural image benchmarks, fine-grained visual recognition, and medical datasets, MUNKEY outperforms all post-hoc baselines. Our results establish that unlearning by design enables fast, deployment-oriented unlearning while preserving predictive performance.

Sources

Related papers