Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models

summary

Video file (mp4)

The gist

The paper presents "a trace-preserving counterfactual that fixes the recorded execution controls while assigning requested identifiers zero contribution" for machine unlearning in large language

In short

The hosts discuss a paper detailing 'state-exact trace-preserving deletion' in large language models. The authors propose a system that allows for bit-for-bit identical removal of data from training sets, proving it works on billion-parameter models like Llama. The discussion focuses on the technical requirements, limitations (like high cost for scattered deletions), and the paper' providing a verifiable standard for data deletion.

Key concepts

Trace-Preserving Deletion
This method allows for exact data removal by recording the entire training process—the 'trace.' When deleting data, the original training plan is replayed, but with deleted examples replaced by zero-loss sequences. This ensures the final model state matches a separate run where those specific examples were physically removed.
State-Exact
This is a strong claim meaning that after deletion, every single parameter of the model and its optimizer state is bit-for-bit identical to what it would be if trained without the deleted data. This goes far beyond simply having the model's behavior change or being statistically similar.
Write-Ahead Log (WAL)
The paper uses a WAL to record every microbatch during training. This log includes fixed records of the seed, learning rate, and optimizer step. It serves as a complete plan that allows researchers to precisely replicate the original training run while implementing data deletions.

Terminology used across episodes

This episode discusses

The paper

Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models · Read on arXiv

Abdullah X

Project AWARE

Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution controls while assigning requested identifiers zero contribution. The guarantee is prospective: the original run must record this execution provenance and retain an eligible uncontaminated checkpoint. Under pinned single-GPU environments, replay from a token store materialized without the requested rows reconstructs a separately executed trace oracle bit-for-bit in model and optimizer state. Pythia 160M is exact across four deletion geometries; Pythia 2.8B matches all 2,775,208,960 model-state elements for a random 5% request; and Llama 3.2 1B is exact after omitting 400 of 4,000 TOFU examples from replay storage. These results establish billion-parameter state exactness. They do not establish cheap deletion, because dispersed requests can force nearly full replay. We release standardized TOFU/OpenUnlearning measurements as descriptive diagnostics only because the frozen campaign lacks the matched controls required for a causal behavioral claim.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models".

Jane: The paper was written by Abdullah X from Project AWARE.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that has a real mouthful of a title: "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models." Jane, I have to say, just reading that title out loud made me want to sit down and pay attention.

Jane: It really does, Tom. And I think the key phrase in there is "state-exact." That's the part that makes this paper different from a lot of the unlearning work we've talked about before. Most of the time, when people say they've "unlearned" something from a model, they mean the model's behavior changed. This paper is saying something much stronger.

Tom: Stronger how? Give me the plain English version.

Jane: So imagine you train a model, and then someone says, "delete these four hundred examples." The usual approach is to try to nudge the model so it acts like it never saw them. This paper instead says: we're going to replay the training, but with those examples physically removed from the data store, and we're going to prove the final model and optimizer state are bit-for-bit identical to a separate run that did exactly that.

Tom: Bit-for-bit. That's not "close enough," that's not "statistically similar." That's every single parameter matching exactly.

Jane: Exactly. And they did it at billion-parameter scale. Pythia 2 point 8B, Llama three point two 1B, all of them. The title is really advertising the core claim: this is exact deletion, not approximate deletion, and it works on models big enough to matter.

Tom: And I love that the title doesn't oversell it. It says "trace-preserving," which is this very careful, precise idea that they're preserving the execution trace of the original training run. They're not claiming they can magically erase knowledge from any model. They're claiming they built a system where, if you set things up right from the start, you can delete data and get an exact state.

Jane: Right. And that distinction is going to come up again and again in this conversation, because it's the difference between a real engineering contribution and a marketing claim. The title is honest about what they're delivering.

Tom: So before we get into the weeds, what's the big picture here? Why should a listener who doesn't work on unlearning care?

Jane: Because data deletion is a legal requirement now. GDPR, right to be forgotten, all of that. And companies have been struggling with what it actually means to "delete" data from a model. This paper gives a concrete, verifiable answer to that question, at least for a specific setup. That's a big deal.

Tom: And it's a big deal that they're being so careful about what they're not claiming. We'll get into that, but for now, let's just say the title sets up a very precise promise, and the rest of the paper is about whether they can keep it.

Summary: Tom: Alright, so we've got the title. Let's talk about what the paper actually does. Jane, you touched on it, but give us the full summary.

Jane: So the core idea is what they call "trace-preserving deletion." The paper defines a counterfactual: what would the model state be if we had trained on the same data, in the same order, with the same random seeds, but with certain examples contributing zero loss? That's the target. And they build a system that can actually reproduce that target exactly.

Tom: And the trick is that they record everything during the original training. They call it a write-ahead log, or WAL. Every microbatch gets a fixed thirty-two-byte record with the seed, the learning rate, the optimizer step, the accumulation boundary. That's the "trace" part.

Jane: Right. So when a deletion request comes in, they don't have to guess what the training looked like. They have the full plan. They just replay it, but with the deleted examples replaced by dummy sequences with zero weight. And the data store they read from physically doesn't contain those rows anymore.

Tom: And the results, Jane. The numbers are wild. Pythia 160M, exact across four different deletion geometries. Pythia 2 point 8B, that's two point seven billion model elements, exact for a random five percent deletion. Llama three point two 1B on TOFU, exact after omitting four hundred of four thousand examples.

Jane: And they're not just checking the model weights. They're checking the optimizer state too. That's important because if you're going to continue training after deletion, the Adam moments matter just as much as the weights. Two models with identical weights but different optimizer states are not the same model going forward.

Tom: That's a really good point. And they're honest about the cost. This is not cheap deletion. If you delete a random five percent of the data, the earliest affected step is basically step zero, so you have to replay almost the entire training run. The paper has a whole section on this.

Jane: Yeah, the locality analysis is actually one of my favorite parts. If you delete a late chunk of data, you can start from a checkpoint right before that chunk and only replay a small suffix. But if the deletions are scattered throughout the whole dataset, you're replaying nearly everything. That's a structural limitation, and they own it.

Tom: So the summary is: they built a system that can do exact deletion, proved it works at billion-parameter scale, and they're very clear about when it's practical and when it's not.

Jane: And they're also very clear about what this doesn't mean. They run the standard OpenUnlearning benchmark on their deleted models, but they explicitly say those results are descriptive, not causal. They don't have the matched controls to claim their method causes better forgetting behavior.

Tom: That kind of restraint is rare. Most papers would have led with the benchmark numbers. They buried them in an appendix and said "don't over-interpret these."

Jane: Exactly. And that's why I trust the state-exactness claim more. When a paper is this careful about what it's not claiming, I'm more inclined to believe what it is claiming.

Tom: So we've got the summary. But I want to dig into what this actually improves. What's the before and after here? What does this paper change about the field?

Improvements: Tom: So we've established what the paper does. But let's talk about what it improves. Jane, what was the state of the art before this?

Jane: Before this, exact unlearning in LLMs mostly meant changing the architecture so that deletion is easy. You shard the model, you isolate data into separate components, you make it so you only have to retrain a small piece. That's the SISA approach, or the more recent S3T and LMEraser work. Those are real contributions, but they change how you train the model in the first place.

Tom: And this paper does something different. It doesn't change the architecture. It changes the execution contract.

Jane: Exactly. They keep a standard transformer, standard AdamW, standard training loop. What they add is provenance: the immutable plan, the write-ahead log, the per-microbatch reseeding. That's the improvement. You can take a conventional training pipeline and make it deletion-exact, as long as you instrument it from the start.

Meng: Can I jump in here? I've been listening and I want to ask about the practical side. You said they keep the standard architecture, but they need this whole provenance infrastructure. What does that actually cost in storage and compute?

Tom: Great question, Meng. The paper has numbers on that. The WAL plus the ID manifest is tiny compared to a checkpoint. For Pythia 2 point 8B, it's about four megabytes of provenance versus eleven gigabytes for one checkpoint. That's zero point zero three seven percent overhead. The compute cost is basically zero during training because you're just writing a few bytes per microbatch.

Meng: Okay, so the ongoing cost is negligible. But the deletion cost, that's the real question. You mentioned replaying almost the entire run for scattered deletions. That's not a small cost.

Jane: Right, and that's the honest trade-off. The paper has this order-statistic calculation that shows if you delete one uniformly random example, you expect to replay about half the run. If you delete two hundred examples, you're at ninety-nine point five percent. So this is not a method for frequent, scattered deletions. It's a method for occasional, localized deletions, or for cases where you absolutely need the exact state.

Meng: So who actually uses this? I'm trying to think about where this fits in a real deployment.

Tom: That's the question, Meng. I think the answer is regulated fine-tuning. If you're a company fine-tuning a model on customer data and you get a deletion request, you need to be able to say "we removed that data" with some confidence. This gives you a verifiable answer. You can produce the hash of the new model state and show it matches the trace-preserving oracle.

Meng: But the cost. If the deletion is scattered, you're redoing the whole fine-tune. That's expensive.

Jane: It is. But the alternative is either not being able to delete at all, or doing an approximate deletion and hoping it's good enough. For some applications, exactness is worth the compute.

Lalam: If I may add a perspective here. The improvement this paper brings is not just technical. It's cultural. It establishes that exactness is a meaningful target for unlearning, not just behavioral similarity. That changes how we evaluate deletion claims. Regulators and auditors can ask for state-level evidence, not just "the model seems to have forgotten." That is a significant shift in what counts as compliance.

Tom: Lalam, that's a really good point. The paper is essentially saying: if you want to claim deletion, here's what a rigorous claim looks like. And that's a contribution that goes beyond any single experiment.

Meng: I still want to know about the failure modes. What happens if the environment changes? Different GPU, different PyTorch version?

Jane: They're upfront about that. The exactness claim is pinned to a specific environment. Same GPU, same software stack, same deterministic settings. If you change any of that, the bit-for-bit guarantee goes away. They're not claiming hardware-independent determinism.

Meng: So it's fragile.

Tom: It's fragile in the sense that it's environment-specific. But it's robust in the sense that within that environment, it's exact. And the paper's whole point is that the environment is part of the specification. You can't separate the deletion claim from the execution context.

Lalam: And that is precisely why this is an improvement. The field has been treating unlearning as a property of the model alone. This paper says it's a property of the model, the data, the execution plan, and the environment together. That is a more mature way to think about the problem.

Conclusion: Tom: Alright, we're wrapping up our discussion of "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models." Jane, give us the final summary.

Jane: So this paper delivers a very specific, very strong result: if you instrument your training run from the start, recording the execution plan and per-microbatch controls, you can replay that run with deleted examples physically removed from the data store and get a model and optimizer state that is bit-for-bit identical to a separately executed trace-preserving oracle. They proved it on Pythia 160M, Pythia 2 point 8B, and Llama three point two 1B on TOFU.

Tom: And they were honest about the limits. It's not cheap for scattered deletions. It requires prospective instrumentation. It's pinned to a specific environment. And the behavioral benchmark results are descriptive, not causal.

Jane: Right. And that honesty is part of why I think this paper matters. It gives the field a clear target: state exactness is achievable, here's how, and here's what it costs. Future work can build on that or argue against it, but they can't ignore it.

Lalam: The broader impact is that this reframes deletion as an engineering problem with verifiable outputs, rather than a vague behavioral aspiration. That has implications for regulation, for auditing, and for how we talk about what a model has or hasn't learned.

Meng: And from a practical standpoint, it gives teams a concrete checklist: record the plan, pin the environment, keep eligible checkpoints, and you can offer exact deletion for at least some request patterns. That's more than most systems can say today.

Tom: Well said, all of you. So we're saying goodbye to this paper. It's a strong contribution, a careful contribution, and one that I suspect will be cited for years as the reference point for what exact deletion actually means.

Jane: Absolutely. And now we're ready to move on to the next paper. Thanks for listening, everyone. We'll see you on the next episode.

Tom: Take care, folks.

More episodes

← Home