SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation

arXiv:2606.00380 · cs.CV, cs.AI · Submitted 2026-05-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation".

Jane: The paper was written by the authors from Loughborough University and Newcastle University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we're cracking open a fresh one from the arXiv — it's called "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation." And Jane, I gotta say, that title is a mouthful, but it's hiding something really cool.

Jane: It really is, Tom. And honestly, the word that jumps out at me is "reproducible." Because if you've ever tried to compare two different machine unlearning papers, you know it's a mess. One team trains for ten epochs, another trains for fifty. One uses one seed, another uses five. You can't tell if Method A is actually better than Method B or if they just got lucky with their random number generator.

Tom: Right, and that's exactly what the authors — Petros Andreou, Jamie Lanyon, Axel Finke, and Georgina Cosma from Loughborough University — are tackling. They've built this framework called SUPREME to standardize the whole evaluation pipeline. And I love that they're not just proposing a new unlearning method; they're building the measuring stick.

Jane: Exactly. And for our listeners who might be new to this, let's just say machine unlearning is the idea of making a trained model "forget" specific data. Like, imagine you trained a face recognition system and you need to remove one person's images because they asked for their data to be deleted. You don't want to retrain the whole thing from scratch — that's expensive — so you want a clever way to scrub that person out.

Tom: And the problem is, how do you know the scrub actually worked? That's where SUPREME comes in. It runs the training, the unlearning, and the evaluation all under the same controlled conditions, across multiple seeds, so you can actually trust the comparison. And the authors are making it open source, which is huge.

Jane: It is. And I love that they've designed it to be extensible. You want to add a new unlearning method? You just implement an interface and register it. You want to test it on a new dataset? Same deal. It's not a locked box; it's a toolkit.

Tom: And the multi-GPU part in the title — that's the part that gets me excited, because running ten seeds of a Vision Transformer on a single GPU could take days. SUPREME spreads that across multiple GPUs, which makes multi-seed evaluation actually feasible for regular research labs.

Jane: Right, and that's the key. Because the paper shows that single-seed results can be wildly misleading. We're going to dig into those numbers in a bit, but trust me, the variance they found is scary.

Tom: Scary and fascinating. So stick around, because next we're going to break down what they actually did and what they found. This is going to be good.

Summary of the Paper: Tom: So Jane, we've set the stage. Now let's get into what SUPREME actually does. The paper lays out a three-stage pipeline: training, unlearning, and evaluation. And the clever part is that every stage is seeded, so every method gets tested under identical starting conditions.

Jane: Right, and that's so important. They use ten training seeds — seeds two hundred sixty through two hundred sixty-nine — and they test two scenarios. The first is full-class unlearning, where you forget all images of five specific celebrities from the Pins Face Recognition dataset. The second is random-sample unlearning, where you randomly remove a tiny zero point one percent slice of the training data.

Tom: And they test this on two very different models: ResNet18, which is a classic convolutional network, and a Vision Transformer, which is the modern attention-based architecture. So they're covering both ends of the spectrum.

Jane: And the results? Oh boy. The variance across seeds is the headline. For example, with Random Labels unlearning on ResNet18 in the random-sample scenario, the forget-set accuracy difference was minus forty-eight point eight nine percentage points, but the standard deviation was plus or minus thirty-four points. That's enormous.

Tom: Yeah, that means one seed might show the method working great, and another seed might show it completely failing. If you only ran one seed, you could publish either result and it would be a coin flip which one you got.

Jane: Exactly. And the same story with Selective Synaptic Dampening on the Vision Transformer — minus fifty-five points with a standard deviation of almost thirty-eight points. These are not small fluctuations; these are game-changing differences.

Tom: And that's the whole point of the paper. They're not just reporting these numbers to show which method is best. They're showing that without multi-seed evaluation, you literally cannot trust a single result.

Jane: And there's another layer to it. They also look at multiple criteria, not just accuracy. They measure forgetting, utility, parametric equivalence, privacy, efficiency, and resources. So a method might look great on accuracy but terrible on privacy, and you'd never know if you only looked at one metric.

Tom: Right, and they show that in the full-class scenario, the membership inference attack scores — that's the privacy metric — barely moved across all methods. But the accuracy differences ranged from zero to eighty-nine percentage points. So if you only looked at privacy, you'd think all methods were equal, which is clearly not the case.

Jane: That's such a good point, Tom. It really shows that you need the full picture. And that's what SUPREME gives you — the full picture, across seeds, across metrics, across models.

Tom: And across GPUs, which is the part that makes it practical. But I want to get into the practical side of this with Meng, because I know he's going to have questions about actually running this thing.

Improvements Suggested by the Paper: Tom: So we've talked about what SUPREME does and what it found. But what does this paper actually improve? What does it change about how we do research?

Jane: I think the biggest improvement is the multi-GPU support. And I want to bring in Meng here, because he's the engineer who actually has to run these things.

Meng: Thanks, Jane. Yeah, so the multi-GPU part is genuinely a big deal. The paper says it's the first image classification unlearning framework to distribute training, unlearning, and evaluation across multiple devices. And they do it through Lightning Fabric, which supports DDP, FSDP, and DeepSpeed ZeRO. That means you can scale from one GPU to many without rewriting your code.

Tom: And that's not just a convenience thing, right? It changes what experiments are even possible.

Meng: Absolutely. Think about it. For each training seed, you need to train the original model, then for each forget class, you need to retrain a baseline model from scratch. That's the expensive part. In their full-class scenario with five classes and ten seeds, you're looking at fifty baseline trainings. On a single GPU, that could take weeks. With multiple GPUs, you can actually finish it in a reasonable time.

Jane: And that's the improvement that makes the whole multi-seed philosophy practical. Without it, people would just keep running single seeds because that's all they can afford.

Meng: Right. And there's another practical detail I like. They have this accelerated model initialization that defers parameter materialization until the model is moved to its device. That sounds technical, but it just means less memory overhead, which matters when you're running many experiments in parallel.

Tom: And they also support mixed precision and tensor cores, which speeds things up even more. So it's not just about scaling out; it's about making each run faster too.

Jane: And I want to bring in Lu here, because I think there's a bigger picture to this. This isn't just about making the current experiments faster.

Lu: Exactly, Jane. The improvement here is that it makes the research question itself more rigorous. By standardizing the evaluation and making multi-seed runs affordable, SUPREME raises the bar for what counts as a valid comparison. A paper that only reports single-seed results will look weak next to one that uses this framework.

Meng: And that's a good thing for the field. It means fewer false claims, fewer wasted replications. But I do want to note one thing: the paper is honest about a limitation. They ran their experiments on a single GPU to maintain exact numerical parity with the reference implementations. So the multi-GPU results might not be bit-for-bit identical to single-GPU runs.

Tom: That's a fair caveat, but it doesn't take away from the fact that this framework is a huge step forward. And I'm curious about what this means for the future of unlearning research. Lu, what do you think?

Lu: I think this opens the door to much larger studies. Instead of ten seeds on one dataset, you could do fifty seeds across multiple datasets. Instead of just ResNet18 and ViT, you could test a whole zoo of architectures. The framework makes that feasible, and that's going to give us much more reliable knowledge about which unlearning methods actually work.

Jane: And that's the hook for our next segment — where does this go from here? What's the bigger impact on the world?

Conclusion: Tom: Alright, we've covered the title, the summary, and the improvements. Now let's wrap this up and talk about what "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation" really means for the world.

Jane: I think the biggest takeaway is that this paper is about trust. When someone claims their unlearning method works, we now have a standard way to check. And the paper shows that without this standard, we might be believing results that are just noise.

Tom: And that's not just an academic problem. Think about the real-world applications. Companies need to comply with data deletion requests. Regulators need to verify that a model has actually forgotten someone's data. If we can't reliably evaluate unlearning, we can't trust those compliance claims.

Lu: And that's where I see the real impact. This framework could become the common benchmark that regulators and companies both use. It's like having a standardized crash test for cars — you can't just claim your car is safe; you have to pass the test.

Meng: And from a practical standpoint, the fact that it's open source and extensible means it can grow with the field. New methods, new metrics, new datasets — they can all be added without rewriting the whole framework.

Jane: And the authors are clear that this is a demonstration, not a definitive ranking. They used one dataset, two models, and ten seeds. But the framework is designed for much bigger studies, and that's what excites me.

Tom: Same here. And I love that they're honest about the limitations. They mention that multi-GPU runs don't exactly reproduce single-device results, and they note that the demonstration is confined to one dataset. That's the kind of scientific humility that makes their claims credible.

Lu: And the future work is exciting too. They want to decompose the across-seed variance into its training, unlearning, and evaluation parts. That would tell us exactly where the randomness comes from, which would help researchers design even better experiments.

Tom: So to sum it up — SUPREME gives us a standardized, scalable, and honest way to evaluate machine unlearning. It shows that single-seed results can be dangerously misleading, and it provides the tools to do better.

Jane: And that's a big deal, because machine unlearning is only going to become more important as data privacy regulations tighten. This paper gives us the foundation to build on.

Tom: Well said, Jane. That's a wrap on "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation." Thanks to everyone for listening, and we'll see you next time with another paper to break down.

Jane: Take care, everyone. And remember — when it comes to unlearning, trust but verify.

Loughborough University · Newcastle University

cs.CV, cs.AI

Submitted: 2026-05-29

Updated: 2026-09-18

Comments: 17 pages. Code available at https://github.com/pedroandreou/supreme-unlearning

Code: https://github.com/pedroandreou/supreme-unlearning

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: The paper introduces SUPREME, an open-source framework for evaluating machine unlearning methods in image classification.

Key concepts

Image Unlearning
This is the idea of making a trained model 'forget' specific data. For example, if you have a face recognition system and need to remove one person's images, unlearning is the process of scrubbing that person out without retraining the entire expensive model from scratch.
Reproducible Evaluation
This means running experiments under controlled conditions so results can be trusted. SUPREME achieves this by running training, unlearning, and evaluation stages under identical starting conditions and across multiple random seeds to check for reliable performance differences.
Multi-GPU Framework
SUPREME is the first image classification unlearning framework to distribute training, unlearning, and evaluation across multiple devices. It uses tools like Lightning Fabric to scale experiments from one GPU to many without requiring code rewrites.
Standardized Metrics
The paper measures multiple criteria beyond just accuracy, including forgetting, utility, parametric equivalence, privacy, efficiency, and resources. This holistic view helps determine if a method is truly effective across all important aspects.

Terminology

Summary

The paper introduces SUPREME, an open-source framework for evaluating machine unlearning methods in image classification. The authors state: "Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. The framework makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds." The framework is available at https://github.com/pedroandreou/supreme-unlearning.

The paper motivates the need for reproducible unlearning evaluation with three factors: "First, the number of proposed methods and metrics is increasing, and comparing them under the same training setup and seeds is required to identify which methods perform best. Second, evaluation now covers larger model architectures such as Vision Transformers and larger datasets, which increases the compute cost of every training run. Third, recent work shows that single-seed results can misrepresent a method's performance, because outcomes depend on initial weights, data ordering, and the stochasticity of the unlearning step [2,11]."

The authors note that existing frameworks include "MUBox [12], which benchmarks methods across several scenarios; ERASURE [4], a modular toolkit for different data domains; and Deep Unlearn [2], which tests methods across multiple random initialisations. Each runs on a single GPU at full precision, and of these only Deep Unlearn supports multi-seed evaluation."

The paper's contributions are: "An extensible framework with a registry-based design covering datasets, model architectures, unlearning methods, evaluation metrics, and unlearning scenarios. New components are added by implementing an interface and registering a module path, without modifying framework code"; "A hardware-agnostic architecture with multi-GPU support built on PyTorch and Lightning Fabric, including DDP, FSDP, and DeepSpeed ZeRO. Distribution applies to training, unlearning, and evaluation — to our knowledge the first image classification unlearning framework to do so; and A demonstration on Pins Face Recognition. We evaluate the integrated methods on ResNet18 and ViT under full-class and random-sample unlearning across ten seeds."

The paper defines: "Let Minit be a model parameterised by initial weights which may be randomly initialised or pre-trained and let Mo be the model trained on some training set D consisting of image–label pairs (x, y). Machine unlearning seeks to remove the influence of some forget set Df ⊆ D from Mo. This is done by applying some unlearning method to Mo to obtain an unlearned model Mu. The goal is that Mu mimics as closely as possible the retrained baseline Mr, i.e., the model which has been trained from scratch on the retain set Dr:= D Df."

SUPREME supports two unlearning scenarios: targeted and random-sample unlearning. These differ in how the set of forget targets C is defined. For targeted unlearning, each c ∈ C is a class (or sub-class) label and the corresponding forget set is Df:= (x, y) ∈ D y = c. For random-sample unlearning, each c ∈ C is a number in (0, 1] and the corresponding forget set Df is a subset – sampled uniformly at random – of D of size ⌈cD⌉. The scenarios also differ in test set definitions: "targeted unlearning uses held-out samples of the forgotten target (Df′:= (x, y) ∈ D′ y = c and Dr′:= D′ Df′) to test whether unlearning generalises beyond the training samples, while random-sample unlearning has no such held-out set and evaluates on the training forget samples directly (Df′:= Df and Dr′:= D′)."

The pipeline has three stages: "Stage 1 trains Mo on D. For each forget target c ∈ C, the framework trains Mr on Dr. For each unlearning method a ∈ A, Stage 2 applies a to Mo to obtain Mu, and Stage 3 evaluates Mu against Mr using the configured metrics."

The pipeline has three properties: Seeded randomness where "Each seed determines initial weights, data ordering, the random forget-set draw (under the random-sample scenario), and the stochastic operations inside the unlearning step. Running every method under the same seed configuration ensures that every a ∈ A is evaluated under identical starting conditions for a given seed, isolating method differences from pipeline randomness"; Distributed execution where "All three stages execute across P devices. Training and unlearning use gradient synchronisation, where gradients are averaged across devices after each backward pass so all devices apply the same parameter update. Evaluation uses result aggregation, where per-batch metric values are collected across devices and averaged into the final metric value; and Registry-based extensibility where Datasets, model architectures, unlearning methods, evaluation metrics, and unlearning scenarios are registered via module paths."

SUPREME is built on PyTorch and Lightning Fabric. The framework supports DDP, FSDP, and DeepSpeed ZeRO-1/2/3 as distributed strategies, selectable at runtime. Two optimisations reduce overhead: enabling tensor cores for matrix operations, and accelerated model initialisation that defers parameter materialisation until the model is moved to its device.

The vision models are: "ResNet18 [10] is a four-block residual convolutional network with 64–512 channels per block, trained from scratch on 32 × 32 inputs. ViT [6] is fine-tuned from google/vit-base-patch16-224, which divides 224 × 224 inputs into 16 × 16 patches and processes them through 12 transformer layers with 12 attention heads and 768 hidden dimensions."

Image preprocessing: "ResNet18 inputs are resized to 32 × 32 with standard training-time augmentation (random crop, horizontal flip, rotation) and per-channel normalisation computed from the training set. ViT inputs are resized to 256 × 256, centre-cropped to 224 × 224, horizontally flipped, and normalised with ImageNet statistics, as required by the pre-trained model. Augmentation is disabled during unlearning and evaluation."

Regarding hardware: "SUPREME's architecture supports multi-GPU execution. Because distributed gradient averaging introduces non-determinism, multi-GPU runs do not exactly reproduce single-device results; the experiments in this paper therefore use a single device."

Unlearning quality is evaluated under seven criteria: behavioural equivalence, parametric equivalence, privacy, forgetting, utility, efficiency, and resources. The metrics include: Layer-wise distance, Activation distance, JS-Divergence, Completeness, Accuracy differences, Loss differences, and Membership Inference Attack differences. Computational cost metrics include Time, Memory Usage, and Compute Utilisation.

The demonstration uses Pins Face Recognition [1], an image classification benchmark of 17,534 facial images across 105 celebrity identities. ResNet18 uses statistics µ = (0.516, 0.419, 0.373) and σ = (0.286, 0.255, 0.246) computed from the training set. ViT uses ImageNet statistics, µ = (0.485, 0.456, 0.406) and σ = (0.229, 0.224, 0.225).

Two unlearning scenarios are evaluated: "Full-class unlearning removes all samples for five identities: alex lawther, bill gates, danielle panabaker, hugh jackman, and josh radnor. Random-sample unlearning removes a 0.1 % subset of training samples drawn from across all classes. The experiments use I = 10 training runs with seeds 260–269... with a single unlearning and evaluation seed per training seed (J = K = 1). At J = K = 1, the seed formulae in Algorithm 1 reduce to su = se = st. The unlearning methods are applied to Mo, with Mr as the baseline. Experiments ran on a single NVIDIA L40S GPU (48 GB VRAM) to maintain exact numerical parity with the reference implementations [3,7,8], following the same choice made by OpenUnlearning [5] in the text domain."

Table 1 reports test-accuracy differences between Mu and Mr on Df′ (∆AccDf′) and Dr′ (∆AccDr′), and the layer-wise weight distance Layer between the two models. Table 4 reports activation distances and membership inference attack (MIA) score differences (∆MIA).

"Several method-and-scenario combinations produced large standard deviations over the ten seeds. The two largest forget-set accuracy (∆AccDf′) standard deviations both occurred in the random-sample scenario. Random Labels on ResNet18 gave ∆AccDf′ = −48.89 ± 34.03 percentage points, and Selective Synaptic Dampening on ViT gave −55.00 ± 37.99. A single seed can therefore differ from the ten-seed mean by tens of percentage points. The observed spread combines four sources of randomness: initial weights, data ordering, the random forget-set draw, and stochastic operations inside the unlearning method."

"In the full-class scenario, the mean MIA difference (∆MIA) in Table 4 stays within ±0.05 for every method on both model architectures, and seven of the twelve satisfy ∆MIA ≤ 0.02. On the same runs, ∆AccDf′ ranges from 0.00 to 89.44 percentage points on ResNet18 and from 0.00 to 35.52 on ViT. ∆MIA therefore does not separate the methods in this setting, whereas ∆AccDf′ does. The methods may have similar privacy behaviour, or the MIA probe may be too weak on Pins Face Recognition."

"We introduced SUPREME, an open-source framework for evaluating image classification unlearning, registry-based across datasets, model architectures, unlearning methods, evaluation metrics, and scenarios, and able to distribute training, unlearning, and evaluation across multiple devices through Lightning Fabric. Our single-GPU demonstration on Pins Face Recognition produced ten-seed results across both model architectures and both scenarios. The demonstration is confined to a single dataset, two architectures, and a single device, so the reported numbers characterise the framework rather than a definitive ranking of the methods; broader conclusions would require the larger-scale, multi-seed studies the framework is designed to enable. Future work will extend the demonstration to the other registered components and use the per-stage seeds to decompose the across-seed spread into its training, unlearning, and evaluation parts."

Petros Andreou is supported by a PhD studentship funded by Darktrace Ltd. The funders had no role in the study design, data collection, analysis, interpretation, or in the decision to submit this work for publication.

"SUPREME is an evaluation framework for machine unlearning methods. It is method-agnostic and introduces no new unlearning method; it standardises how existing methods are compared. The demonstration uses Pins Face Recognition, a publicly available academic benchmark of celebrity images. Identity is used as the classification target only to test whether identity-level information can be removed from a trained model. We do not propose or endorse face recognition as a deployed application... SUPREME is a technical evaluation framework. Low forget-set accuracy or small parametric distance under our metrics does not constitute legal verification of data erasure under any specific regulation."

The framework supports datasets including CIFAR-10, CIFAR-20, CIFAR-100, PinsFaceRecognition, Caltech-101; models ResNet18, Vision Transformer (ViT); unlearning methods Retrain, Fine-Tuning (FT), Bad Teacher (BadT), Random Labels (RL), Unlearning by Selective Impair and Repair (UNSIR), Selective Synaptic Dampening (SSD), Loss-Free Selective Synaptic Dampening (LFSSD); unlearning scenarios Targeted (full-class, sub-class), Random sample; accelerators CPU, GPU, MPS, TPU; precision modes including "64-bit Double Precision (64-true), 32-bit Full Precision (32-true), 16-bit Mixed Precision (16-mixed), 16-bit Brain Floating Point Mixed Precision (bf16-mixed), 16-bit True Precision (16-true), 16-bit True Brain Floating Point Precision (bf16-true), 8-bit Floating Point via NVIDIA TransformerEngine with bfloat16 weights (transformer-engine), 8-bit Floating Point via NVIDIA TransformerEngine with float16 weights (transformer-engine-float16), 8-bit Integer Inference via BitsandBytes (int8), 8-bit Integer Fine-Tuning via BitsandBytes (int8-training), 4-bit NormalFloat via BitsandBytes (nf4), 4-bit NormalFloat with Double Quantisation via BitsandBytes (nf4-dq), 4-bit Floating Point via BitsandBytes (fp4), 4-bit Floating Point with Double Quantisation via BitsandBytes (fp4-dq); and distributed strategies Distributed Data Parallel (DDP), Fully Sharded Data Parallel (FSDP), Zero Redundancy Optimiser (DeepSpeed ZeRO) Stage-1/2/3."

Improvements for AI systems

Based on the paper, here are specific improvements I can make to AI systems:

Improvement: Build a distributed training/unlearning/evaluation pipeline using PyTorch Lightning Fabric that supports DDP, FSDP, and DeepSpeed ZeRO-1/2/3 strategies, with gradient synchronization across devices during training/unlearning and all-gather result aggregation during evaluation.

What it can do: Evaluate machine unlearning methods on large models (e.g., ViT) and datasets across multiple GPUs simultaneously, reducing wall-clock time for multi-seed experiments from days to hours. For example, running 10 seeds of ResNet18 full-class unlearning on Pins Face Recognition (5 forget classes) would previously require sequential single-GPU runs; now it can be parallelized across P devices with near-linear speedup.

Improvement: Implement a pipeline with three independent seed types—training seed (st), unlearning seed (su), and evaluation seed (se)—that can be varied independently or jointly. The system should support J unlearning seeds per training seed and K evaluation seeds per unlearning seed, with deterministic random number generation at each stage.

Improvement: Create a modular architecture with registries for datasets (CIFAR-10/100, Pins Face Recognition, Caltech-101), models (ResNet18, ViT), unlearning methods (FT, BadT, UNSIR, RL, SSD, LFSSD), and metrics (7 criteria, 20+ metrics). New components are added by implementing an interface and registering a module path, without modifying core framework code.

Improvement: Implement a comprehensive evaluation suite covering seven criteria: behavioural equivalence (activation distance, JS-divergence, completeness), parametric equivalence (layer-wise distance), privacy (MIA), forgetting (forget-set accuracy/loss/completeness/activation/JS), utility (retain-set metrics), efficiency (time), and resources (memory, compute utilization). All metrics are computed as differences between the unlearned model and the retrained baseline.

Improvement: Implement two distinct unlearning scenarios with different data-split logic: (a) targeted unlearning where forget sets are label-defined (e.g., all images of bill gates) and test splits use held-out samples; (b) random-sample unlearning where forget sets are random subsets (e.g., 0.1% of training data) and evaluation uses the training forget samples directly with the full test set as retain.

Improvement: Support 14 precision modes (from 64-bit double to 4-bit quantized) and 4 accelerator types (CPU, GPU, MPS, TPU), with runtime selection of distributed strategy and precision. Include tensor-core optimization and deferred parameter materialization for faster multi-seed runs.

Improvement: Automatically compute and report mean ± standard deviation across all seeds for every metric, with per-seed raw values stored. Flag methods where the standard deviation exceeds a configurable threshold (e.g., >10 percentage points) as highly variable and suggest increasing seed count.

Abstract

Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. SUPREME makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds. The framework is available at https://github.com/pedroandreou/supreme-unlearning.

Related papers