OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets

arXiv:2607.13037 · cs.AI · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets".

Jane: The paper was written by Haolin Xue from Northwestern Polytechnical University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: We’ve seen the core problem—the struggle to pinpoint specific author contributions in massive AI datasets—and we’re ready to talk about what this whole project is called: OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets. What does that name tell us about the fundamental shift in how we approach data management?

Jane: The title suggests a level of detail that moves way beyond just knowing which file a piece of training data came from, right? It promises to give us granular accountability by linking everything back to the specific individuals who created it.

Lu: It implies that we aren’re moving past the era of simply asking "what source is this" toward a more precise inquiry like "who contributed this sentence," which is exactly what's necessary for complex AI training sets.

Meng: That shift from broad file-level tracing to targeted record-level provenance indicates that the system is designed to manage complexity, not just one single source, making it practical for real-world scale.

Lalam: The cultural significance of being able to pinpoint authorship is immense; it honors the individual contributor by ensuring their work isn't lost in a sea of data scrubbing or mass deletion.

Tom: But we’ also have to factor in the token level, which Jane noted earlier, because that’s where the AI actually reads. This title suggests they are tracking provenance down to that smallest possible unit of input, which is truly groundbreaking.

Jane: Yes, and that's vital because when we can trace back from the token—the absolute smallest piece of data—we can finally track influence with a precision that was simply impossible before.

Lu: This gives us a verifiable path from the abstract idea of "unlearning" to a concrete, actionable step by connecting directly to those individual records in the data set.

Meng: It’s not just tracking; the system is engineered for practical integration, meaning it's built to be plugged into existing pipelines without requiring massive, manual rewrites.

Lalam: We are developing a framework that allows us to respect individual contribution while preserving the collective power of knowledge in our digital age.

Summary: Tom: Now, having established the concept, let’s look at how OriginBlame actually functions internally. Can you walk us through its core mechanics as described in the paper?

Jane: It uses a clever three-layer architecture—Authors, Sections, and Document-Index—to maintain a highly organized structure for tracking everything. Imagine every single row of data has pointers back to the specific source files and their creators.

Lu: The fundamental idea is propagating that author identity across transformations; as the original text gets tokenized and packed into binary shards, the links remain completely intact throughout every `ob.track` call in a processing pipeline.

Meng: It’s a content-addressable system that uses cryptographic hashes to ensure everything is immutable and verifiable, which means we can never lose track of what has changed or where it originated from.

Lalam: It creates a complete, unbroken history for every single piece of data, ensuring that the story of human creativity remains intact even as the data is transformed into its final form.

Tom: And Jane mentioned earlier that this isn’t just about tracking; we need to handle co-authorship where one author might contribute to a section that another author then assemble. How does the architecture manage this?

Jane: That’s where the Sections layer becomes essential, allowing multiple authors to be linked to a single source file path, so it can accurately represent complex co-authorship without needing to split the content itself.

Lu: The system tracks that the entire lineage of these sections down to individual records within those files is necessary for providing a clear and auditable chain for compliance.

Meng: It handles this complexity by using hashes and tracking which specific section hashes contribute to each document-index entry, making it incredibly robust against manual input errors.

Lalam: This architecture ensures that the contribution of every person is recorded accurately, regardless of how many other people were working on the same project.

Improvements: Tom: That’s a robust system, but what truly sets OriginBlame apart is its ability to solve the critical problem of over-deletion when someone requests their data be removed, right?

Jane: Absolutely, Tom. The study found that without this record-level approach, we would have to delete massive amounts of irrelevant data—up to one hundred one times more than necessary—just to remove one author’s work.

Lu: The improvement lies in the fact that record-level provenance allows us to surgically target only those specific lines where a particular author’s contribution exists, making it precise.

Meng: Quantitatively, this translates into reducing that over-deletion factor dramatically, down to about one point three times the volume required for targeted removal, which is a massive practical win for data hygiene and sustainability.

Lalam: The impact here is much more ethical; we are not punishing other contributors or wasting resources by deleting material that doesn't belong to the person who made the request.

Tom: And it’s not just about deletion; what happens when the data changes, like when you edit a training file? How does OriginBlame handle that challenge?

Jane: It has this clever two-phase reconcile mechanism, where it first checks if the content hashes match, and if they don't, it uses semantic matching to see if the meaning is similar.

Lu: This dual approach—hash matching followed by embedding similarity—means that even when a data record is slightly altered or mutated, we can still trace its connection back to its original source and maintain continuity.

Meng: The results show that this system recovers between ninety-six percent and ninety-eight percent of the original provenance links even after the data has been edited, which is incredibly reliable for projects that are constantly evolving.

Lalam: This resilience means our ability to manage or unlearn data remains consistent over time, allowing us to trust the history of our AI models more deeply.

Conclusion: Tom: We've discussed the core problem and the brilliant solution through record and token-level provenance, as well as the impressive practical improvements in reducing waste.

Jane: It’s a real relief to see such a robust system that solves this complex governance issue for model trainers, providing clarity where there was previously only confusion.

Lu: The theoretical significance of this work is that we are now capable of truly granular control over our data lineage, opening up new and very exciting possibilities for future research methods.

Meng: The low overhead and verifiable architecture suggest a highly practical tool that can be integrated into large-scale systems without requiring a complete overhaul.

Lalam: I believe the cultural impact is profound—we are building more ethical and accountable AI systems by respecting individual data rights at the token level.

Tom: That sense valuing contribution is key; it’s about acknowledging that the data isn't just a monolithic pile but a collection of specific human efforts.

Jane: And Lu, you mentioned structure—the three-layer architecture really seems to be what makes this possible, locking down authorship as the primary point of reference throughout the various transformation steps.

Lu: Exactly, Jane; it creates a verifiable path from the raw source all the way through to tokenization that is incredibly robust and dependable for future generations.

Meng: I’m interested in how this translates into real-world compliance, ensuring that these precise records can be easily queried by regulatory bodies without being overly complex.

Lalam: We are moving toward a system where data governance is inherently transparent, which guarantees fairness across different communities and contributors.

Tom: It's a massive leap from simply having to throw away entire datasets when compliance demands it, right?

Jane: This work shows that precision is not only possible but that the results are significantly better for machine unlearning tasks compared to random approaches.

Lu: I am excited to see how this opens up new lines of research in the future.

Meng: I just hope we see these tools put into production immediately, making it a practical tool for large-scale data management.

Lalam: A final thought on OriginBlame: Record- and Token-Level Data Provenance for AI Training Datasets is that its precision will change how we define data ownership forever.

Haolin Xue

Northwestern Polytechnical University

cs.AI

Submitted: 2026-08-22

Updated: 2026-08-25

Code: https://github.com/tzbkk/originblame

Importance score: 92/100

The gist: The paper presents OriginBlame (ob), a record- and token-level data provenance system designed to solve the gap between knowing who to forget and knowing which data to forget in machine unlearning

Key concepts

Record- and Token-Level Provenance
This refers to tracking the origin of data not just by file, but down to the specific record (row) or even the smallest unit of text (token). This granular detail allows for precise attribution and management within massive AI datasets.
Three-Layer Architecture
OriginBlame uses a structured system involving Authors, Sections, and Document-Index layers. This organization maintains a highly detailed pointer structure, allowing the system to accurately track complex co-authorship across different sources.
Data Provenance
The concept of tracking data lineage—the complete history of where data came from and how it was transformed. OriginBlame uses this to ensure an unbroken, verifiable path for every piece of data throughout the AI training process.

Terminology

Summary

The paper presents OriginBlame (ob), a record- and token-level data provenance system designed to solve the gap between knowing who to forget and knowing which data to forget in machine unlearning scenarios.

When a data contributor requests removal, model trainers face a practical gap because existing tools cannot locate which training records belong to a given author at the required granularity. Existing provenance systems operate at file or dataset level, forcing catastrophic over-deletion. This issue is highlighted by benchmarks like TOFU and MUSE, which evaluate unlearning on synthetically curated forget sets but lack a mechanism to locate affected data at training granularity. Traditional tools (DVC, LakeFS, Delta Lake) operate only at the file or dataset level; experiment management tools track metadata but not data origins.

OriginBlame addresses this by propagating author identity through the data processing pipeline and resolves revocation requests into precise forget sets via a three-layer architecture:

  1. Authors Layer: Stores identities and revocation tags (Author ID, name, email, revoked boolean). The revoked field is the single source of truth for revocation.

  2. Sections Layer: Stores file-level copyright (Section hash, path, authors list, license).

  3. Document-Index Layer: Stores per-record provenance (Line hash, file name, sources list linking to section hashes).

This architecture is supplemented by an independent Token-Index Layer.

The system is designed with the following principles:

  • Precision: It locates specific data lines, not files or datasets.

  • Verifiability: Provenance must be independently auditable.

  • Minimal Invasiveness: Integration requires only a few lines of code—a single ob.track call per record.

The core workflow involves the ob.track(data, file, section, embedding, model) function:

  1. It computes the SHA-256 hash of the data (JSON serialization).

  2. It resolves the active source list into section hashes.

  3. It checks for duplicates (idempotent).

  4. It writes to process-isolated files using Write-Ahead Logging (WAL) protection, ensuring data recovery if the process crashes before ob clean merges the files.

Query Functionality:

  • ob blame: Performs a forward query, tracing the provenance chain from line hash to sources (section hashes) to sections (author IDs).

  • ob show: Performs a reverse query, traversing authors to sections to document-index.

Operation Log: The .ob/log file records state-changing operations (registration, revocation, purge), providing an audit trail.

ob supports three granularity levels of revocation:

  1. Author-level revoke: Sets revoked=true on the author record. This tag cascades lazily through the chain to find all affected sections and document-index entries, resulting in a list of every data line traceable to the revoked author.

  2. Section-level revoke: Toggles revoked=true on a specific section record, affecting only document-index entries referencing that section without revoking the author.

  3. Line-level revoke: Toggles revoked=true on a single document-index entry, enabling granular line-level withdrawal.

The Purge operation physically deletes revoked data lines from the specified file, with a safety constraint ensuring no purge occurs without prior revocation.

To support tokenization and packing into binary shards, ob utilizes the Token-Index layer:

  • Each entry stores token count, sources (linking to the author chain), tokenizer, and a `revoked boolean.

  • This layer is position-indexed, allowing the i-th entry’s token range in the packed binary to be reconstructed from cumulative sums of preceding entries’ values.

  • ob generate-set --tokenizer gpt2-o forget.bin produces a binary bitmask where each bit corresponds to one token-index entry (1 = revoked, 0 = active), which is directly usable by unlearning algorithms.

Over-deletion Reduction:

Evaluation on 219,555 Wikipedia pages demonstrates that record-level provenance eliminates dataset-level over-deletion (from 101 times to 1.3 times). This is consistent across all tested scales, where the largest share author had only 1.3–1.6 times over-deletion, while the smallest share authors faced about 100 times under dataset-level revocation.

Performance and Scalability:

  • Integration overhead is modest: HuggingFace adds 2.0% throughput and Datatrove adds 13.8% throughput. The storage overhead is low, ranging from 1.33 times to 1.23 times.

  • Query latency is fast: ob show completes in 80 ms at 220k lines, and ob revoke completes in 22 ms.

Machine Unlearning Effectiveness:

The paper validates that fine-grained provenance improves unlearning efficacy. Using Qwen3-1.7B on zhwiki data, provenance-based forget sets substantially outperform random baselines, achieving a 42% improvement in forgetting compared to random baselines using the NPO algorithm.

OriginBlame has four limitations:

  1. Incremental adoption is difficult; users cannot retroactively add provenance to existing datasets.

  2. The parser ecosystem is immature (only a MediaWiki parser is currently available).

  3. It maintains single-hop provenance, tracking only direct mappings from raw data to final output.

  4. The index must be rebuilt when data changes, and indexed variants do not consistently outperform the full-scan path due to parallelism amortization.

Improvements for AI systems

Based on a meticulous review of the OriginBlame (ob) framework, the following improvements can be implemented in AI training pipelines. These enhancements transform data governance from a coarse, destructive process into a precise, verifiable engineering discipline.

The integration of ob allows for the construction of an entire Provenance-Aware Training Stack. This stack is characterized by the seamless propagation of author identity and data origins through all processing stages, ensuring traceability from raw source material to final tokenized model input.

Specific Improvements:

  1. Granular Data Provenance (Line and Token Level):
  • Instead of relying on file-level attribution (which is inadequate for collaborative content), ob maintains a three-layer, content-addressable hierarchy (Author to Section to Document-Index).

  • The Token-Index Layer extends this record to the token level, allowing us to map specific tokens in the final packed binary output back to their exact source file and author. This is critical for LLM training where data is tokenized and packed before the source metadata is lost.

  1. Deterministic Revocation via Lazy Cascading:
  • The system utilizes a tagging model (e.g., revoked = true on the Author record). Revocation does not trigger immediate bulk deletion. Instead, queries (show, blame) lazily cascade this status down the chain (Author to Section to Document-Index).

  • Result: This mechanism allows for precise, targeted data removal, eliminating the catastrophic over-deletion (up to 101 times) that plagues current dataset-level compliance efforts.

  1. Resilience and Data Integrity via Reconcile:

The system incorporates a two-phase reconciliation strategy. When source files are edited or mutated, the provenance links break. ob uses:

  • Phase 1 (Hash Matching): Identifies unchanged content using cryptographic hashes (SHA-256).

  • Phase 2 (Semantic Matching): Uses embedding similarity to recover provenance for lines that have been rewritten or slightly altered.

This ensures the integrity of the traceability chain is maintained even when source data is dynamic, preventing the loss of audit trails.

By integrating this provenance system, the resulting AI training pipeline gains four critical capabilities:

1. High-Fidelity Compliance and Ethical Data Management:

  • Function: The system can execute automated Right to Erasure requests by generating a precise, verifiable Forget Set that contains only the data lines attributed to a specific individual (e.g., Author X).

  • Improvement: This moves AI compliance from an expensive, destructive guess to an efficient, targeted engineering task.

2. Superior Unlearning Efficacy:

  • The ob framework allows for the creation of author-specific forget sets. When used with unlearning algorithms (like NPO or RMU), this level of precision significantly outperforms random sampling.

  • Improvement: The resulting model achieves a 42% improvement in forgetting the targeted data, while simultaneously demonstrating superior utility preservation compared to random baselines.

3. Debugging and Bias Attribution:

  • The ability to query the Document-Index allows researchers to trace any observed model behavior (e.g, a specific bias or failure mode) back through the token stream to identify the exact source data points (lines/tokens) that contributed most heavily to that outcome.

  • Improvement: Drastically reduces debugging time and complexity by allowing for precise attribution of training data influence.

4. Operational Efficiency and Scalability:

  • ob is designed with an O(1) lookup capability via its sharding and Bucket-Routing Index. This allows for rapid querying (e.g., show --author NAME) even at massive scales (about 220 k pages).

  • Improvement: The system operates with low latency, maintaining a negligible throughput overhead (<4%) on modern data processing frameworks (HuggingFace/Datatrove), ensuring that the operational cost of compliance does not impede training speed.

Sources

Related papers