Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text

arXiv:2608.16868 · cs.CL, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Towards Computational Provenance".

Jane: The paper investigates whether generated text can carry verifiable evidence about which causally relevant internal state occurred during model computation,

Tom: First, who's behind it and why it matters.

Title and authors: Jane: We’ve just talked about what this paper is trying to prove—that we can track which internal computation happened even when the answer is identical. Let’s start by looking at the title, "Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text."

Tom: It sounds a bit heavy, but it gets right to the core problem: how do we get proof about what happened inside the AI without needing to see every single calculation?

Lu: The authors are Benjamin Belay and their team. They are looking at computational provenance, which is basically tracking the history of computation in these models.

Meng: So they’re not trying to map out every single neuron firing, but rather identifying a causally relevant internal state and seeing if we can carry evidence of that state into the text.

Tom: That’s right. It moves away from trying to reconstruct the entire internal reasoning process, which is practically impossible for large models.

Jane: They are focusing on a narrower question: can we identify a causally relevant internal state, verify which state occurred, and make that state determine a detectable signal in the model’s generated output?

Lu: They test this concept in two controlled architectures: a modular feed-forward neural network and a transformer-based model.

Tom: So they’re showing it works across different styles of AI, not just one specific design, which makes it feel more generalizable.

Jane: The paper suggests that this approach could complement existing methods for interpreting AI by providing verifiable evidence about the internal computation that produced the output.

The paper's summary: Tom: We’ve established the concept of computational provenance, but what exactly is this system doing? It’s not just a fancy label; it’s a mechanism to link an output back to a specific internal decision point.

Jane: Think about it this way: if you ask an AI question and get one answer, and then you change one tiny thing internally that *should* lead to the same answer, but the model outputs slightly different wording, that's where this paper is looking for evidence.

Lu: They set up a specific arithmetic pathway where different internal paths can produce the same final answer through two discrete intermediate states, z2 and z3.

Meng: They deliberately switch between these paths and then authenticate the state that was actually used in each execution to see if it leaves a statistical pattern in the text.

Tom: So they are essentially saying, "Okay, this word choice is linked to internal state A or internal state B," even though the final number is identical.

Jane: The paper shows that this connection can be made verifiable through cryptographic receipts and then used to select a specific subset of wording choices.

Lu: They have a template for the generated text where each group has eight alternatives, and they designate four as favored and four as unfavored based on the verified state.

Meng: So, if state z2 was 'four', only certain words get picked in that sentence structure, but the meaning of the sentence stays exactly what it is.

Tom: It’s about preserving evidence about a causally relevant part of the computation while keeping the rest of the output unchanged for an evaluator.

The paper's improvements: Jane: Now that we know how they did it, what are the suggested improvements in "Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text"? It’s not just a one-off trick.

Tom: One major improvement they suggest is combining three different forms of evidence to make the system really robust.

Lu: They argue that you need intervention tests to show the recorded state actually participates in the computation, not just acting as an unrelated label.

Meng: Then you need authenticated receipts to establish precisely which state was observed during a particular execution.

Jane: And finally, you need that statistical signal in the generated output to preserve evidence of that verified state.

Tom: That combination is key; it shows the required causal pathway reproduced across five feed-forward models and three transformers, which is quite impressive consistency.

Lu: Plus, they found that separately trained models achieved one hundred twenty-eight out of one hundred twenty-eight on both their public and protected end-to-end evaluations <ref:2608.16868#pg0,protected end-to-end evaluations>.

Meng: So this isn't just a lab result; it’s showing how this provenance mechanism can be transferred between different model architectures.

Conclusion: Tom: To wrap things up, the main implication of "Towards Computational Provenance: Carrying Causal-State Evidence in Generated Text" is that we can start building a framework for oversight that goes beyond just checking the final answer.

Jane: It suggests a new verification framework for larger models that checks evidence associated with verified internal computation rather than relying only on the model’s own explanation.

Lu: It’s a controlled proof of concept in this specific arithmetic task, but it lays out a path for larger models where we can check evidence associated with verified internal computation.

Meng: This capability supports oversight by providing evidence that selected causally relevant parts of the internal computation were connected to an observable output, which is crucial when human supervisors can’t directly evaluate difficult artifacts.

Tom: It moves us toward systems where we don't just trust the answer; we have a way to verify *why* that answer was generated in a specific way.

Jane: The authors are showing that by combining intervention tests, receipts, and signal preservation, we can distinguish between executions that follow different internal paths but would otherwise appear identical.

Lu: It’s a controlled proof of concept in a finite arithmetic task with an explicitly constructed discrete pathway.

Meng: This is really about making the internal reasoning more transparent in a way that is practical for real-world AI deployment, not just theoretical research.

Benjamin Belay

cs.CL, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-17

Comments: 16 pages, 1 figure, 7 tables

Journal ref: NeurIPS 2026 Workshop on Interpretability as a Science

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 90/100

The gist: The paper investigates whether generated text can carry verifiable evidence about which causally relevant internal state occurred during model computation, addressing a critical gap in understanding

Key concepts

Causal-State Evidence
This refers to verifiable proof of which specific internal step or state (like z2 or z3) was actually used by the model during its calculation. The paper shows this evidence can be recorded cryptographically and then linked to the final generated text, even if the final answer is the same.
Intervention Tests
These tests involve deliberately changing an internal state (e.g., modifying z2) while keeping the final result consistent. This proves that a specific intermediate state is causally relevant to the computation rather than just being an arbitrary label, confirming its role in determining the output.
State-Dependent Statistical Signal
This is how information about a specific internal state (like z2=4) influences the wording of the generated text. For each possible state value, certain words are favored; by analyzing which words appear most frequently, researchers can infer which internal state was used.
Authenticated Receipts
These are cryptographic records that verify two things: first, that a specific intermediate state event occurred during an execution, and second, the actual value of that state. These receipts serve as the bridge between the model's hidden computation and the observable text.

Terminology

Summary

The paper investigates whether generated text can carry verifiable evidence about which causally relevant internal state occurred during model computation, addressing a critical gap in understanding how internal reasoning connects to observable outputs. The gist: evidence of a verified, causally relevant internal state can be preserved in generated text even when the final answer is unchanged.

How it works

The study tests this concept using two controlled architectures: a modular feed-forward neural network and a transformer-based model, both trained on the same arithmetic task where different internal paths can produce the same final answer. The core mechanism involves deliberately switching between these paths and authenticating the state actually used to determine a statistical pattern in the generated text, which is then tested by a detector <ref:2608.16868#pg3>.

The controlled arithmetic pathway specifies that for an input prompt of four numbers, the computation must pass through two discrete intermediate states, z2 and z3: z1 = (a + b) mod 16, z2 = (z1 + c) mod 16, z3 = (5z2 + d) mod 16. (1) Both architectures are built so that the computation must pass through two explicit, discrete states, z2 and z3 <ref:2608.16868#pg5>. To create matched executions that yield the same answer through different paths, an intervention replaces z2 with a modified value, such as z′2 = (z2 + 8) mod 16 <ref:2608.16868#pg5>. The final modulo-8 operation ensures that despite the change in internal states, the observed answer remains identical: z′3 mod 8 = z3 mod 8 <ref:2608.16868#pg5>.

How it works

To distinguish which internal path was taken, the researchers implemented a system for recording and verifying these intermediate states using cryptographic records called receipts <ref:2608.16868#pg6>. They use two kinds of state evidence: An exact receipt identifies the particular intermediate-state event from one execution <ref:2608.16868#pg6>, and An abstract receipt records the corresponding state value, such as z2 = 4, which is the identity used to select the later state-specific statistical signal <ref:2608.16868#pg6>. Only after these checks succeed is the authenticated value of z2 allowed to determine the statistical signal used during text generation <ref:2608.16868#pg6>.

How it works

The causal information from the verified state is carried into the generated text by introducing variations in wording within a fixed-length output sentence <ref:2608.16868#pg6>. The text follows a template where Each group contains eight permitted alternatives <ref:2608.16868#pg6>. For each verified value of z2, four of the eight alternatives in every group are designated as favoured and the other four as unfavoured <ref:2608.16868#pg6>. This state-dependent subset influences word choice, allowing different internal states to leave different statistical patterns in the wording without changing what the text reports <ref:2608.16868#pg6>.

How it works

A detector then compares these observed word choices against patterns associated with all 16 possible values of z2 <ref:2608.16868#pg6>. The detector gives higher scores when more of the observed words match those favoured by a candidate state <ref:2608.16868#pg6>. A state is accepted only when its score exceeds a threshold fixed on separate calibration data and is higher than all other candidate states <ref:2608.16868#pg6>. This process successfully identifies which state-specific wording pattern is present while the meaning of the generated text remains unchanged <ref:2608.16868#pg6>.

How it works

The experimental design evaluates three parts: whether the model uses the intermediate state causally, whether that state can be verified, and whether its associated signal can be recovered from the generated text <ref:2608.16868#pg8>. The results show that the required causal pathway reproduced across five feed-forward models and three transformers <ref:2608.16868#pg8>, and separately trained models achieved 128/128 on both their public and protected end-to-end evaluations <ref:2608.16868#pg8>.

How it works

The conclusion establishes that the result depends on "combining three forms of evidence: Intervention tests show that the recorded state participates in the computation rather than acting as an unrelated label; authenticated receipts establish which state was actually observed during a particular execution; and the statistical signal preserves evidence of that verified state in the generated output <ref:2608.16868#pg10>. This suggests a new verification framework for larger models that would check evidence associated with verified internal computation rather than rely only on the model’s own explanation <ref:2608.16868#pg10>. The work remains a controlled proof of concept in a finite arithmetic task with an explicitly constructed discrete pathway" <ref:2608.16868#pg10>.

REFERENCES

Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamile˙ Lukosiˇ ut¯ e, Amanda Askell, Andy Jones, Anna Chen, et al. Measuring progress on scalable ˙oversight for large language models. arXiv preprint arXiv:2211.03540, 2022. doi: 10.48550/arXiv:2211.03540.

Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. Weakto-strong generalization: Eliciting strong capabilities with weak supervision. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pp. 4971–5012, 2024. URL https://proceedings.mlr.press/v235/burns24b.html.

Coalition for Content Provenance and Authenticity. C2PA technical specification, version 2.2, 2025. URL https://spec.c2pa.org/specifications/specifications/2.2/specs/C2PA Specification.html.

Sumanth Dathathri, Abigail See, Sumedh Ghaisas, Po-Sen Huang, Rob McAdam, Johannes Welbl, Vandana Bachani, Alex Kaskasoli, Robert Stanforth, Tatiana Matejovicova, Jamie Hayes, Nidhi Vyas, Majd Al Merey, Jonah Brown-Cohen. Scalable watermarking for identifying large language model outputs. Nature 634:818–823, 2024. doi: 10.1038/s41586-024-08025-4.

Kit Fraser-Taliente, Subhash Kantamneni, Euan Ong, Dan Mossing, Christina Lu, Paul C. Bogdan, Emmanuel Ameisen, James Chen, Dzmitry Kishylau, Adam Pearce. Natural language autoencoders produce unsupervised explanations of LLM activations. 2026. URL https://transformer-circuits.pub/2026/nla/.

Atticus Geiger, Hanson Lu, Thomas Icard, and Christopher Potts. Causal abstractions of neural networks. In Advances in Neural Information Processing Systems, volume 34, pp. 9574–9586, 2021. URL https://proceedings.neurips.cc/paper/2021/hash/4f5c422f4d49a5a807eda27434231040-Abstract.

Atticus Geiger, Zhengxuan Wu, Hanson Lu, Joshua Rozner, Elisa Kreiss, Thomas Icard, Noah D. Goodman, and Christopher Potts. Inducing causal structure for interpretable neural networks. In Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 7324–7338, 2022. URL https://proceedings.mlr.press/v162/geiger22a.html.

Zahra Ghodsi, Tianyu Gu, and Siddharth Garg.

Improvements for AI systems

  1. Bold header: Direct verification of causally relevant internal states in generated text. The system can identify a causally relevant internal state, verify which state occurred, and make that state determine a detectable signal in the model’s generated output, allowing for verifiable evidence about how an output was produced, rather than relying only on the answer or the model’s own explanation.

  2. Bold header: Distinguishing answer-equivalent executions through subtle wording patterns. The improved system can differentiate between two executions that produce the same final answer, semantic content, and sampling randomness by detecting a subtle statistical pattern in the generated text tied to a verified, causally relevant internal state, even when the final answer is unchanged.

  3. Bold header: Enabling scalable oversight for complex model outputs. This capability supports oversight by providing evidence that selected causally relevant parts of its internal computation were connected to an observable output, which is crucial when human supervisors cannot directly evaluate difficult artifacts, effectively complementing methods like sparse autoencoders and Natural Language Autoencoders.

  4. Bold header: Robustness across different model architectures. The system's provenance mechanism can be transferred between a modular feed-forward neural network and a transformer-based model, ensuring that the evidence of internal state is preserved regardless of the underlying architecture used for generation.

Sources

Related papers