Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models".
Jane: The paper was written by Abdullah X from Project AWARE.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back to the show, everyone. Today we're digging into a paper that has a real mouthful of a title: "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models." Jane, I have to say, just reading that title out loud made me want to sit down and pay attention.
Jane: It really does, Tom. And I think the key phrase in there is "state-exact." That's the part that makes this paper different from a lot of the unlearning work we've talked about before. Most of the time, when people say they've "unlearned" something from a model, they mean the model's behavior changed. This paper is saying something much stronger.
Tom: Stronger how? Give me the plain English version.
Jane: So imagine you train a model, and then someone says, "delete these four hundred examples." The usual approach is to try to nudge the model so it acts like it never saw them. This paper instead says: we're going to replay the training, but with those examples physically removed from the data store, and we're going to prove the final model and optimizer state are bit-for-bit identical to a separate run that did exactly that.
Tom: Bit-for-bit. That's not "close enough," that's not "statistically similar." That's every single parameter matching exactly.
Jane: Exactly. And they did it at billion-parameter scale. Pythia 2 point 8B, Llama three point two 1B, all of them. The title is really advertising the core claim: this is exact deletion, not approximate deletion, and it works on models big enough to matter.
Tom: And I love that the title doesn't oversell it. It says "trace-preserving," which is this very careful, precise idea that they're preserving the execution trace of the original training run. They're not claiming they can magically erase knowledge from any model. They're claiming they built a system where, if you set things up right from the start, you can delete data and get an exact state.
Jane: Right. And that distinction is going to come up again and again in this conversation, because it's the difference between a real engineering contribution and a marketing claim. The title is honest about what they're delivering.
Tom: So before we get into the weeds, what's the big picture here? Why should a listener who doesn't work on unlearning care?
Jane: Because data deletion is a legal requirement now. GDPR, right to be forgotten, all of that. And companies have been struggling with what it actually means to "delete" data from a model. This paper gives a concrete, verifiable answer to that question, at least for a specific setup. That's a big deal.
Tom: And it's a big deal that they're being so careful about what they're not claiming. We'll get into that, but for now, let's just say the title sets up a very precise promise, and the rest of the paper is about whether they can keep it.
Summary: Tom: Alright, so we've got the title. Let's talk about what the paper actually does. Jane, you touched on it, but give us the full summary.
Jane: So the core idea is what they call "trace-preserving deletion." The paper defines a counterfactual: what would the model state be if we had trained on the same data, in the same order, with the same random seeds, but with certain examples contributing zero loss? That's the target. And they build a system that can actually reproduce that target exactly.
Tom: And the trick is that they record everything during the original training. They call it a write-ahead log, or WAL. Every microbatch gets a fixed thirty-two-byte record with the seed, the learning rate, the optimizer step, the accumulation boundary. That's the "trace" part.
Jane: Right. So when a deletion request comes in, they don't have to guess what the training looked like. They have the full plan. They just replay it, but with the deleted examples replaced by dummy sequences with zero weight. And the data store they read from physically doesn't contain those rows anymore.
Tom: And the results, Jane. The numbers are wild. Pythia 160M, exact across four different deletion geometries. Pythia 2 point 8B, that's two point seven billion model elements, exact for a random five percent deletion. Llama three point two 1B on TOFU, exact after omitting four hundred of four thousand examples.
Jane: And they're not just checking the model weights. They're checking the optimizer state too. That's important because if you're going to continue training after deletion, the Adam moments matter just as much as the weights. Two models with identical weights but different optimizer states are not the same model going forward.
Tom: That's a really good point. And they're honest about the cost. This is not cheap deletion. If you delete a random five percent of the data, the earliest affected step is basically step zero, so you have to replay almost the entire training run. The paper has a whole section on this.
Jane: Yeah, the locality analysis is actually one of my favorite parts. If you delete a late chunk of data, you can start from a checkpoint right before that chunk and only replay a small suffix. But if the deletions are scattered throughout the whole dataset, you're replaying nearly everything. That's a structural limitation, and they own it.
Tom: So the summary is: they built a system that can do exact deletion, proved it works at billion-parameter scale, and they're very clear about when it's practical and when it's not.
Jane: And they're also very clear about what this doesn't mean. They run the standard OpenUnlearning benchmark on their deleted models, but they explicitly say those results are descriptive, not causal. They don't have the matched controls to claim their method causes better forgetting behavior.
Tom: That kind of restraint is rare. Most papers would have led with the benchmark numbers. They buried them in an appendix and said "don't over-interpret these."
Jane: Exactly. And that's why I trust the state-exactness claim more. When a paper is this careful about what it's not claiming, I'm more inclined to believe what it is claiming.
Tom: So we've got the summary. But I want to dig into what this actually improves. What's the before and after here? What does this paper change about the field?
Improvements: Tom: So we've established what the paper does. But let's talk about what it improves. Jane, what was the state of the art before this?
Jane: Before this, exact unlearning in LLMs mostly meant changing the architecture so that deletion is easy. You shard the model, you isolate data into separate components, you make it so you only have to retrain a small piece. That's the SISA approach, or the more recent S3T and LMEraser work. Those are real contributions, but they change how you train the model in the first place.
Tom: And this paper does something different. It doesn't change the architecture. It changes the execution contract.
Jane: Exactly. They keep a standard transformer, standard AdamW, standard training loop. What they add is provenance: the immutable plan, the write-ahead log, the per-microbatch reseeding. That's the improvement. You can take a conventional training pipeline and make it deletion-exact, as long as you instrument it from the start.
Meng: Can I jump in here? I've been listening and I want to ask about the practical side. You said they keep the standard architecture, but they need this whole provenance infrastructure. What does that actually cost in storage and compute?
Tom: Great question, Meng. The paper has numbers on that. The WAL plus the ID manifest is tiny compared to a checkpoint. For Pythia 2 point 8B, it's about four megabytes of provenance versus eleven gigabytes for one checkpoint. That's zero point zero three seven percent overhead. The compute cost is basically zero during training because you're just writing a few bytes per microbatch.
Meng: Okay, so the ongoing cost is negligible. But the deletion cost, that's the real question. You mentioned replaying almost the entire run for scattered deletions. That's not a small cost.
Jane: Right, and that's the honest trade-off. The paper has this order-statistic calculation that shows if you delete one uniformly random example, you expect to replay about half the run. If you delete two hundred examples, you're at ninety-nine point five percent. So this is not a method for frequent, scattered deletions. It's a method for occasional, localized deletions, or for cases where you absolutely need the exact state.
Meng: So who actually uses this? I'm trying to think about where this fits in a real deployment.
Tom: That's the question, Meng. I think the answer is regulated fine-tuning. If you're a company fine-tuning a model on customer data and you get a deletion request, you need to be able to say "we removed that data" with some confidence. This gives you a verifiable answer. You can produce the hash of the new model state and show it matches the trace-preserving oracle.
Meng: But the cost. If the deletion is scattered, you're redoing the whole fine-tune. That's expensive.
Jane: It is. But the alternative is either not being able to delete at all, or doing an approximate deletion and hoping it's good enough. For some applications, exactness is worth the compute.
Lalam: If I may add a perspective here. The improvement this paper brings is not just technical. It's cultural. It establishes that exactness is a meaningful target for unlearning, not just behavioral similarity. That changes how we evaluate deletion claims. Regulators and auditors can ask for state-level evidence, not just "the model seems to have forgotten." That is a significant shift in what counts as compliance.
Tom: Lalam, that's a really good point. The paper is essentially saying: if you want to claim deletion, here's what a rigorous claim looks like. And that's a contribution that goes beyond any single experiment.
Meng: I still want to know about the failure modes. What happens if the environment changes? Different GPU, different PyTorch version?
Jane: They're upfront about that. The exactness claim is pinned to a specific environment. Same GPU, same software stack, same deterministic settings. If you change any of that, the bit-for-bit guarantee goes away. They're not claiming hardware-independent determinism.
Meng: So it's fragile.
Tom: It's fragile in the sense that it's environment-specific. But it's robust in the sense that within that environment, it's exact. And the paper's whole point is that the environment is part of the specification. You can't separate the deletion claim from the execution context.
Lalam: And that is precisely why this is an improvement. The field has been treating unlearning as a property of the model alone. This paper says it's a property of the model, the data, the execution plan, and the environment together. That is a more mature way to think about the problem.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Unlearning at Scale: State-Exact Trace-Preserving Deletion in Billion-Parameter Language Models." Jane, give us the final summary.
Jane: So this paper delivers a very specific, very strong result: if you instrument your training run from the start, recording the execution plan and per-microbatch controls, you can replay that run with deleted examples physically removed from the data store and get a model and optimizer state that is bit-for-bit identical to a separately executed trace-preserving oracle. They proved it on Pythia 160M, Pythia 2 point 8B, and Llama three point two 1B on TOFU.
Tom: And they were honest about the limits. It's not cheap for scattered deletions. It requires prospective instrumentation. It's pinned to a specific environment. And the behavioral benchmark results are descriptive, not causal.
Jane: Right. And that honesty is part of why I think this paper matters. It gives the field a clear target: state exactness is achievable, here's how, and here's what it costs. Future work can build on that or argue against it, but they can't ignore it.
Lalam: The broader impact is that this reframes deletion as an engineering problem with verifiable outputs, rather than a vague behavioral aspiration. That has implications for regulation, for auditing, and for how we talk about what a model has or hasn't learned.
Meng: And from a practical standpoint, it gives teams a concrete checklist: record the plan, pin the environment, keep eligible checkpoints, and you can offer exact deletion for at least some request patterns. That's more than most systems can say today.
Tom: Well said, all of you. So we're saying goodbye to this paper. It's a strong contribution, a careful contribution, and one that I suspect will be cited for years as the reference point for what exact deletion actually means.
Jane: Absolutely. And now we're ready to move on to the next paper. Thanks for listening, everyone. We'll see you on the next episode.
Tom: Take care, folks.
Abdullah X
Project AWARE
cs.LG, cs.AI, cs.CR
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 18 pages, 2 figures, 7 tables. Code and frozen research records available at https://github.com/abdullah-x-bd/unlearning (v0.3.1)
Code: https://github.com/abdullah-x-bd/unlearning
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 58/100
The gist: The paper presents "a trace-preserving counterfactual that fixes the recorded execution controls while assigning requested identifiers zero contribution" for machine unlearning in large language
Key concepts
- Trace-Preserving Deletion
- This method allows for exact data removal by recording the entire training process—the 'trace.' When deleting data, the original training plan is replayed, but with deleted examples replaced by zero-loss sequences. This ensures the final model state matches a separate run where those specific examples were physically removed.
- State-Exact
- This is a strong claim meaning that after deletion, every single parameter of the model and its optimizer state is bit-for-bit identical to what it would be if trained without the deleted data. This goes far beyond simply having the model's behavior change or being statistically similar.
- Write-Ahead Log (WAL)
- The paper uses a WAL to record every microbatch during training. This log includes fixed records of the seed, learning rate, and optimizer step. It serves as a complete plan that allows researchers to precisely replicate the original training run while implementing data deletions.
Terminology
Summary
The paper presents a trace-preserving counterfactual that fixes the recorded execution controls while assigning requested identifiers zero contribution
for machine unlearning in large language models. The authors define a prospectively instrumented training continuation
that can reproduce a deletion counterfactual exactly after selected examples leave its replay dataset.
The guarantee is prospective: the original run must record this execution provenance and retain an eligible uncontaminated checkpoint.
The paper formalizes training as an executed program. Let D = ((i, zi))ni=1 be an indexed tokenized record store and let F be a set of requested identifiers.
The replayed model-and-optimizer state is written as St = (θt, omegat),
where θt is the model state and omegat is optimizer state.
The training execution is represented by an ordered plan P = (p1,..., pM), pj = (Ij, sj, ηj, uj, aj),
where "Ij is the ordered tuple of identifiers occupying microbatch plan record j, sj is its random seed, ηj is the learning-rate value, uj is the logical optimizer step to which the microbatch contributes, and aj ∈ 0, 1 denotes the end of a gradient-accumulation segment."
The paper defines trace-preserving deletion (Definition 3.1): "For a deletion request F, define TD(P, F) as the execution that retains every microbatch plan record pj but maps each identifier in Ij ∩ F to a fixed dummy example whose loss weight is zero. Retained identifiers, seeds, learning rates, logical optimizer-step indices, accumulation boundaries, and batch shapes are unchanged. At an accumulation boundary, an optimizer transition is executed if and only if at least one retained weighted contribution occurred in that segment; otherwise the logical boundary is preserved but the optimizer transition is skipped."
Trace exactness (Definition 3.2): A deletion procedure is trace-exact for (P, F, E) if its terminal state is bit-identical, in the recorded state dtypes, to S−F,traceT for both model and optimizer state.
Physical replay redaction (Definition 3.3): "A replay store is physically redacted for F if no token row, label row, or metadata record corresponding to an identifier in F is present in that store. The immutable execution plan may retain those identifiers solely as logical slot markers."
The paper distinguishes this from dense retained-data retraining: "A conventional retained-data run can remove all identifiers in F and densely repack the remaining identifiers into new microbatches. Let RP(P, F) denote a new plan built from those retained examples. In general, RP(P, F) ≠ TD(P, F)."
The system uses an immutable JSONL execution plan from the ordered sample identifiers and the training configuration.
Per-microbatch seeds are deterministically derived from a base seed and microbatch index using SHA-256; learning-rate values are quantized to float32 when the plan is created.
A compact binary WAL and an ordered-ID manifest
records training information. The binary WAL record is exactly 32 bytes
containing: Ordered-ID tag (8 bytes), Microbatch seed (8 bytes), Learning rate (4 bytes), Logical optimizer step (4 bytes), Microbatch length (2 bytes), Flags (2 bytes), CRC32 (4 bytes).
The deterministic execution contract includes: "The released experiments fix the Python, NumPy, PyTorch, and CUDA seeds. They disable cuDNN benchmarking and TF32, request deterministic cuDNN behavior and strict PyTorch deterministic algorithms, set CUBLAS WORKSPACE CONFIG=:4096:8, and reseed the RNGs at every microbatch."
The deletion oracle process: "the runner finds the earliest logical optimizer step touched by the forget set and selects the latest eligible checkpoint whose label is no greater than that step... The trace oracle is then executed as a separate run from that checkpoint using the original plan with forgotten slots mapped to zero-weight dummies."
Proposition 4.4 (Trace-exact replay after physical redaction): Under Propositions 4.1 to 4.3, replay from a physically redacted store produces a model-and-optimizer state bit-identical to the separately executed trace-preserving oracle.
The proof proceeds by induction: "Both executions begin at the same checkpoint state. Consider plan record pj. For a retained identifier, oracle and replay read the same frozen token and label arrays. For a forgotten identifier, both executions substitute the same fixed dummy arrays and weight zero before any forgotten-row lookup; therefore physical absence of the forgotten row does not change the numerical microbatch."
Pythia 160M (Biderman et al., 2023) from the frozen step143000 revision and a fixed-length WikiText103-derived causal-LM token store
on "one NVIDIA RTX A6000, microbatch size 4, gradient accumulation 4, one epoch, AdamW with learning rate 10−5, weight decay 0.01, 5% warmup, cosine decay, and checkpoints every 250 logical optimizer steps. The execution plan contains 5,064 microbatches."
Four forget-set geometries were tested: early 0.1%, middle 1%, late 1%, and random 5%.
Three deletion policies were compared: "slot mask is the trace-preserving physical-redaction replay. filter removes forgotten rows from the forward microbatch, changing batch shape. repacked constructs a new densely packed retained-data plan. Only slot mask is expected to equal the trace oracle."
Results: slot-preserving replay is exact for every tested deletion geometry, while both physical filtering and dense repacking diverge from the trace oracle.
The L2 differences for filter were 3.1943 (early), 1.7244 (middle), 0.3264 (late), 3.2358 (random); for repacked were 3.2366, 2.6036, 1.9461, 3.1935 respectively.
The 2.8B endpoint uses the Pythia 2.8B step143000 revision on one NVIDIA A100-SXM4-80GB... microbatch size is 1 and gradient accumulation is 16. The run contains 20,256 WAL records and 1,266 logical optimizer updates.
Only identity replay and random-5% deletion were tested. "Identity replay is exact across 388 tensors and all 2,775,208,960 model elements. The random-5% deletion removes 1,013 of 20,256 records, and the redacted replay exactly matches the oracle with zero unequal tensors, zero unequal elements, maximum absolute difference 0, L2 difference 0, and an equal optimizer hash."
"Llama 3.2 1B Instruct (Meta AI, 2024) and TOFU (Maini et al., 2024)... The forget10 split identifies 400 dataset examples. Five epochs yield 20,000 sample presentations, represented by 2,500 microbatch plan/WAL records... microbatch size 8, gradient accumulation 4, 625 logical optimizer updates, AdamW peak learning rate 10−7, 20% warmup, and cosine decay. The frozen training log shows
the cumulative mean supervised-token loss decreasing from 2.693 after 50 updates to 2.132 at update 625."
"Identity replay is exact across 147 tensors and 1,498,482,688 model elements. The TOFU deletion removes 400 of 4,000 records; the replay store contains 3,600 retained records... The redacted replay again has zero unequal tensors/elements and an optimizer hash equal to the oracle."
The paper reports: "the Pythia 2.8B random-5% replay takes 3,136.20 seconds versus 3,182.83 seconds for the original continuation, and the Llama/TOFU replay takes 1,715.13 seconds versus 1,722.50 seconds. Both restart from step zero. The method therefore demonstrates billion-parameter state exactness, not sublinear or inexpensive deletion."
The structural limitation is quantified: "Consider a one-pass execution with T distinct presentation positions and m affected positions sampled uniformly without replacement. If Kmin is the earliest affected position, the standard discrete order statistic gives E[Kmin] = (T + 1)/(m + 1). With an ideal checkpoint immediately before that position, the expected fraction of the execution that must be replayed is E[(T − Kmin + 1)/T] = (T + 1)m/(T(m + 1)) ≈ m/(m + 1). Thus
one uniformly random affected position implies roughly 50% expected replay, 20 imply 95.2%, 203 imply 99.5%, and 1,013 imply 99.9%."
"Pythia 160M stores 5,064 such records (162,048 bytes) plus a 1,134,250-byte ordered-ID manifest; Pythia 2.8B stores 20,256 records (648,192 bytes) plus a 3,453,690-byte manifest; Llama/TOFU stores 2,500 records (80,000 bytes) plus an 878,890-byte manifest. The ratio of WAL+manifest to one base checkpoint is
0.200% (Pythia 160M), 0.037% (Pythia 2.8B), 0.019% (Llama 3.2 1B + TOFU)."
The paper explicitly states: "The frozen OpenUnlearning run is not used as evidence that trace-preserving deletion causes effective behavioral forgetting. Two controls required for that interpretation are absent: the official retain90 checkpoint is not a matched local retained-data run, and no pre-fine-tuning base-model TOFU evaluation was frozen."
Descriptive results: "For the 400 per-example forget-set Truth Ratio scores, the original and trace-deletion distributions are nearly coincident in this frozen run: their direct two-sample KS statistic is D = 0.015. Against the independently trained official retain90 reference, D = 0.2075 for the original target and D = 0.2100 for the trace-deletion state, with the corresponding upstream Forget Quality p-values 5.97 × 10−8 and 3.91 × 10−8. Model Utility was
0.354289 (original), 0.354136 (exact deletion), 0.592732 (official retain90)."
The paper lists explicit limitations: All exactness claims are from pinned single-GPU runs... We have not demonstrated multi-GPU exactness.
We claim bitwise equality only for the recorded software, hardware, dtype, attention implementation, and deterministic-kernel settings.
One frozen seed per scientific endpoint.
Behavioral evidence is descriptive only.
No publication-authoritative approximate baseline comparison.
One-shot deletion only.
Prospective continuation and fine-tuning scope.
Checkpoint storage and replay cost.
Non-adversarial provenance integrity.
Limits of replay-store redaction.
The paper concludes: "The result is a systems construction with a narrow operating envelope. It establishes neither universal preference for this counterfactual nor cheap deletion. Randomly dispersed requests can force nearly full replay, and the method requires prospective provenance. We therefore keep standardized OpenUnlearning measurements as secondary descriptive diagnostics and draw no behavioral-efficacy conclusion from them. Replay-store redaction and state exactness do not by themselves establish behavioral forgetting. Connecting these properties requires matched experiments beyond the frozen campaign."
Improvements for AI systems
Based on the scientific paper, here are the specific improvements I can implement in AI systems and what the improved systems can do:
Improvement: Add a replay-aware training contract to AI training pipelines that records:
-
A 32-byte fixed-width binary WAL per microbatch (containing seed, learning rate, optimizer step, accumulation boundary, CRC32)
-
An ordered-ID manifest with SHA-256 integrity hashes
-
Per-microbatch RNG reseeding (Python, NumPy, PyTorch, CUDA)
-
Sum-reduced weighted token loss (not mean-reduced) to preserve numerical semantics under deletion
What the improved system can do: After a deletion request, it can reconstruct a bit-exact model and optimizer state that matches a separately executed trace-preserving oracle, without requiring the deleted data rows to be present in the replay token store.
Abstract
Can a prospectively instrumented training continuation reproduce a deletion counterfactual exactly after selected examples leave its replay dataset? We study a trace-preserving counterfactual that fixes recorded execution controls while assigning requested identifiers zero contribution. The guarantee is prospective: the original run must record this execution provenance and retain an eligible uncontaminated checkpoint. Under pinned single-GPU environments, replay from a token store materialized without the requested rows reconstructs a separately executed trace oracle bit-for-bit in model and optimizer state. Pythia 160M is exact across four deletion geometries; Pythia 2.8B matches all 2,775,208,960 model-state elements for a random 5% request; and Llama 3.2 1B is exact after omitting 400 of 4,000 TOFU examples from replay storage. These results establish billion-parameter state exactness. They do not establish cheap deletion, because dispersed requests can force nearly full replay. We release standardized TOFU/OpenUnlearning measurements as descriptive diagnostics only because the frozen campaign lacks the matched controls required for a causal behavioral claim.
Sources
- Machine Unlearning
- Quantifying Memorization Across Neural Language Models
- Simplicity Prevails: Rethinking Negative Preference Optimization for LLM Unlearning
- Pointer Sentinel Mixture Models
- Unlearning Isn't Deletion: Investigating Reversibility of Machine Unlearning in LLMs
- Position: The Term "Machine Unlearning" Is Overused in LLMs
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks