SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation

summary

Video file (mp4)

The gist

The paper introduces SUPREME, an open-source framework for evaluating machine unlearning methods in image classification.

In short

The episode discusses SUPREME, a multi-GPU framework for reproducible evaluation of image unlearning methods. The hosts explain how SUPREME standardizes training, unlearning, and evaluation across multiple seeds and models to ensure trustworthy comparisons. They conclude that this framework is crucial for building trust in unlearning research as data privacy regulations increase.

Key concepts

Image Unlearning
This is the idea of making a trained model 'forget' specific data. For example, if you have a face recognition system and need to remove one person's images, unlearning is the process of scrubbing that person out without retraining the entire expensive model from scratch.
Reproducible Evaluation
This means running experiments under controlled conditions so results can be trusted. SUPREME achieves this by running training, unlearning, and evaluation stages under identical starting conditions and across multiple random seeds to check for reliable performance differences.
Multi-GPU Framework
SUPREME is the first image classification unlearning framework to distribute training, unlearning, and evaluation across multiple devices. It uses tools like Lightning Fabric to scale experiments from one GPU to many without requiring code rewrites.
Standardized Metrics
The paper measures multiple criteria beyond just accuracy, including forgetting, utility, parametric equivalence, privacy, efficiency, and resources. This holistic view helps determine if a method is truly effective across all important aspects.

Terminology used across episodes

This episode discusses

The paper

SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation · Read on arXiv

Loughborough University · Newcastle University

Machine unlearning removes the influence of specific training data from a trained model without retraining it from scratch. Evaluating an unlearning method requires repeating training, unlearning, and evaluation across multiple seeds, which is computationally expensive. To our knowledge, existing image classification unlearning frameworks run on a single GPU, which limits how many seeds can be evaluated in reasonable time. We introduce SUPREME, an open-source framework that distributes these stages across multiple GPUs. SUPREME makes three contributions: a registry-based design for adding new methods, metrics, models, and scenarios; a multi-GPU architecture supporting multiple accelerators and precision modes; and a demonstration on Pins Face Recognition using ResNet18 and ViT under full-class and random-sample unlearning across ten seeds. The framework is available at https://github.com/pedroandreou/supreme-unlearning.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation".

Jane: The paper was written by the authors from Loughborough University and Newcastle University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone. Today we're cracking open a fresh one from the arXiv — it's called "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation." And Jane, I gotta say, that title is a mouthful, but it's hiding something really cool.

Jane: It really is, Tom. And honestly, the word that jumps out at me is "reproducible." Because if you've ever tried to compare two different machine unlearning papers, you know it's a mess. One team trains for ten epochs, another trains for fifty. One uses one seed, another uses five. You can't tell if Method A is actually better than Method B or if they just got lucky with their random number generator.

Tom: Right, and that's exactly what the authors — Petros Andreou, Jamie Lanyon, Axel Finke, and Georgina Cosma from Loughborough University — are tackling. They've built this framework called SUPREME to standardize the whole evaluation pipeline. And I love that they're not just proposing a new unlearning method; they're building the measuring stick.

Jane: Exactly. And for our listeners who might be new to this, let's just say machine unlearning is the idea of making a trained model "forget" specific data. Like, imagine you trained a face recognition system and you need to remove one person's images because they asked for their data to be deleted. You don't want to retrain the whole thing from scratch — that's expensive — so you want a clever way to scrub that person out.

Tom: And the problem is, how do you know the scrub actually worked? That's where SUPREME comes in. It runs the training, the unlearning, and the evaluation all under the same controlled conditions, across multiple seeds, so you can actually trust the comparison. And the authors are making it open source, which is huge.

Jane: It is. And I love that they've designed it to be extensible. You want to add a new unlearning method? You just implement an interface and register it. You want to test it on a new dataset? Same deal. It's not a locked box; it's a toolkit.

Tom: And the multi-GPU part in the title — that's the part that gets me excited, because running ten seeds of a Vision Transformer on a single GPU could take days. SUPREME spreads that across multiple GPUs, which makes multi-seed evaluation actually feasible for regular research labs.

Jane: Right, and that's the key. Because the paper shows that single-seed results can be wildly misleading. We're going to dig into those numbers in a bit, but trust me, the variance they found is scary.

Tom: Scary and fascinating. So stick around, because next we're going to break down what they actually did and what they found. This is going to be good.

Summary of the Paper: Tom: So Jane, we've set the stage. Now let's get into what SUPREME actually does. The paper lays out a three-stage pipeline: training, unlearning, and evaluation. And the clever part is that every stage is seeded, so every method gets tested under identical starting conditions.

Jane: Right, and that's so important. They use ten training seeds — seeds two hundred sixty through two hundred sixty-nine — and they test two scenarios. The first is full-class unlearning, where you forget all images of five specific celebrities from the Pins Face Recognition dataset. The second is random-sample unlearning, where you randomly remove a tiny zero point one percent slice of the training data.

Tom: And they test this on two very different models: ResNet18, which is a classic convolutional network, and a Vision Transformer, which is the modern attention-based architecture. So they're covering both ends of the spectrum.

Jane: And the results? Oh boy. The variance across seeds is the headline. For example, with Random Labels unlearning on ResNet18 in the random-sample scenario, the forget-set accuracy difference was minus forty-eight point eight nine percentage points, but the standard deviation was plus or minus thirty-four points. That's enormous.

Tom: Yeah, that means one seed might show the method working great, and another seed might show it completely failing. If you only ran one seed, you could publish either result and it would be a coin flip which one you got.

Jane: Exactly. And the same story with Selective Synaptic Dampening on the Vision Transformer — minus fifty-five points with a standard deviation of almost thirty-eight points. These are not small fluctuations; these are game-changing differences.

Tom: And that's the whole point of the paper. They're not just reporting these numbers to show which method is best. They're showing that without multi-seed evaluation, you literally cannot trust a single result.

Jane: And there's another layer to it. They also look at multiple criteria, not just accuracy. They measure forgetting, utility, parametric equivalence, privacy, efficiency, and resources. So a method might look great on accuracy but terrible on privacy, and you'd never know if you only looked at one metric.

Tom: Right, and they show that in the full-class scenario, the membership inference attack scores — that's the privacy metric — barely moved across all methods. But the accuracy differences ranged from zero to eighty-nine percentage points. So if you only looked at privacy, you'd think all methods were equal, which is clearly not the case.

Jane: That's such a good point, Tom. It really shows that you need the full picture. And that's what SUPREME gives you — the full picture, across seeds, across metrics, across models.

Tom: And across GPUs, which is the part that makes it practical. But I want to get into the practical side of this with Meng, because I know he's going to have questions about actually running this thing.

Improvements Suggested by the Paper: Tom: So we've talked about what SUPREME does and what it found. But what does this paper actually improve? What does it change about how we do research?

Jane: I think the biggest improvement is the multi-GPU support. And I want to bring in Meng here, because he's the engineer who actually has to run these things.

Meng: Thanks, Jane. Yeah, so the multi-GPU part is genuinely a big deal. The paper says it's the first image classification unlearning framework to distribute training, unlearning, and evaluation across multiple devices. And they do it through Lightning Fabric, which supports DDP, FSDP, and DeepSpeed ZeRO. That means you can scale from one GPU to many without rewriting your code.

Tom: And that's not just a convenience thing, right? It changes what experiments are even possible.

Meng: Absolutely. Think about it. For each training seed, you need to train the original model, then for each forget class, you need to retrain a baseline model from scratch. That's the expensive part. In their full-class scenario with five classes and ten seeds, you're looking at fifty baseline trainings. On a single GPU, that could take weeks. With multiple GPUs, you can actually finish it in a reasonable time.

Jane: And that's the improvement that makes the whole multi-seed philosophy practical. Without it, people would just keep running single seeds because that's all they can afford.

Meng: Right. And there's another practical detail I like. They have this accelerated model initialization that defers parameter materialization until the model is moved to its device. That sounds technical, but it just means less memory overhead, which matters when you're running many experiments in parallel.

Tom: And they also support mixed precision and tensor cores, which speeds things up even more. So it's not just about scaling out; it's about making each run faster too.

Jane: And I want to bring in Lu here, because I think there's a bigger picture to this. This isn't just about making the current experiments faster.

Lu: Exactly, Jane. The improvement here is that it makes the research question itself more rigorous. By standardizing the evaluation and making multi-seed runs affordable, SUPREME raises the bar for what counts as a valid comparison. A paper that only reports single-seed results will look weak next to one that uses this framework.

Meng: And that's a good thing for the field. It means fewer false claims, fewer wasted replications. But I do want to note one thing: the paper is honest about a limitation. They ran their experiments on a single GPU to maintain exact numerical parity with the reference implementations. So the multi-GPU results might not be bit-for-bit identical to single-GPU runs.

Tom: That's a fair caveat, but it doesn't take away from the fact that this framework is a huge step forward. And I'm curious about what this means for the future of unlearning research. Lu, what do you think?

Lu: I think this opens the door to much larger studies. Instead of ten seeds on one dataset, you could do fifty seeds across multiple datasets. Instead of just ResNet18 and ViT, you could test a whole zoo of architectures. The framework makes that feasible, and that's going to give us much more reliable knowledge about which unlearning methods actually work.

Jane: And that's the hook for our next segment — where does this go from here? What's the bigger impact on the world?

Conclusion: Tom: Alright, we've covered the title, the summary, and the improvements. Now let's wrap this up and talk about what "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation" really means for the world.

Jane: I think the biggest takeaway is that this paper is about trust. When someone claims their unlearning method works, we now have a standard way to check. And the paper shows that without this standard, we might be believing results that are just noise.

Tom: And that's not just an academic problem. Think about the real-world applications. Companies need to comply with data deletion requests. Regulators need to verify that a model has actually forgotten someone's data. If we can't reliably evaluate unlearning, we can't trust those compliance claims.

Lu: And that's where I see the real impact. This framework could become the common benchmark that regulators and companies both use. It's like having a standardized crash test for cars — you can't just claim your car is safe; you have to pass the test.

Meng: And from a practical standpoint, the fact that it's open source and extensible means it can grow with the field. New methods, new metrics, new datasets — they can all be added without rewriting the whole framework.

Jane: And the authors are clear that this is a demonstration, not a definitive ranking. They used one dataset, two models, and ten seeds. But the framework is designed for much bigger studies, and that's what excites me.

Tom: Same here. And I love that they're honest about the limitations. They mention that multi-GPU runs don't exactly reproduce single-device results, and they note that the demonstration is confined to one dataset. That's the kind of scientific humility that makes their claims credible.

Lu: And the future work is exciting too. They want to decompose the across-seed variance into its training, unlearning, and evaluation parts. That would tell us exactly where the randomness comes from, which would help researchers design even better experiments.

Tom: So to sum it up — SUPREME gives us a standardized, scalable, and honest way to evaluate machine unlearning. It shows that single-seed results can be dangerously misleading, and it provides the tools to do better.

Jane: And that's a big deal, because machine unlearning is only going to become more important as data privacy regulations tighten. This paper gives us the foundation to build on.

Tom: Well said, Jane. That's a wrap on "SUPREME: A Multi-GPU Framework for Reproducible Image Unlearning Method Evaluation." Thanks to everyone for listening, and we'll see you next time with another paper to break down.

Jane: Take care, everyone. And remember — when it comes to unlearning, trust but verify.

More episodes

← Home