Correct Is Not Governed: Provenance Integrity in Agentic Workflows

arXiv:2608.12761 · cs.AI, cs.CR · Submitted 2026-08-13 · Read on arXiv

Jesus Salas

Microsoft

cs.AI, cs.CR

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: 19 pages, 2 figures

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: Core Thesis: The paper defines "governed execution" as distinct from task correctness.

Terminology

Summary

Core Thesis: The paper defines governed execution as distinct from task correctness. Correct execution is not necessarily governed execution. A workflow is governed only when the authority behind its decisions, the evidence behind its completion, and the effects of later change are explicit and independently checkable. The authors argue that Task correctness and governed execution are different properties.

The Problem: Agentic workflows are typically evaluated by whether they reach the correct outcome, which is insufficient in institutional settings. "An agent can choose the expected action while relying on a draft, expired waiver, or adjacent-domain exception. It can announce that a task is complete without producing the evidence required by the task's definition of done. It can also reach a current business outcome after a policy change while leaving no account of which prior work became stale, which evidence remained valid, or why only some tasks were repeated. Ordinary success metrics can score each workflow as correct."

The Proposed Solution — Matrix: The paper presents Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Matrix separates deterministic institutional mechanics from model-dependent cognition. Models interpret documents, judge applicability, propose plans, and perform work, while The framework owns stable identity, direct references, transition validity, materialization, declared preconditions, receipt matching, dependency traversal, bounded retry counters, and replay.

Three Forms of Provenance Integrity:

  1. Decision provenance identifies the current authority and accepted facts that materially governed an action, while distinguishing them from merely retrieved or semantically related evidence.

  2. Execution provenance connects each obligation to independently checkable completion evidence and prevents an agent's own completion claim from serving as sufficient proof.

  3. Change provenance records which decisions, tasks, receipts, and prior verdicts depend on a superseded authority or fact so that recovery can invalidate stale work without discarding unaffected proof.

Key Results:

Study 1 (Decision provenance): In Experiment 5, "Every retrieval condition selected the expected action in all 15 repeated runs. Direct RAG was therefore perfect on action accuracy, yet it cited non-governing records in six runs. The model-signaled governed path cited or promoted none. The paper notes: The six direct-RAG errors were stable by scenario: in every seed it cited an internal-tool exception and public-brand standard for the harmless theme task, and a non-authoritative research note for the unresolved biometric task."

Study 2 (Execution provenance): In C4 and C5, The shared Phi plan was exact after generic self-review in every run. The direct condition nevertheless accepted an unsupported DPA completion claim and self-closed 3/3 in both rungs. Matrix retains the invalid first receipt, issues the specific evidence correction, validates the countersigned retry, and closes from the new current receipt set. The paper emphasizes: "The generic self-review control matters: the observed difference is not against a one-shot model that was never asked to check its plan. The self-review produced a complete plan, but it did not create an independent execution verifier."

Study 3 (Change provenance): In C6, Matrix traversed declared dependencies and reissued DPA once per seed. The direct control conservatively repeated all six tasks per seed. The result: a sixfold difference in task reexecution on one declared amendment (18 tasks reissued by full rerun versus 3 by dependency-scoped recovery).

Study 4 (Structural refusal): "Phi proposed escalation on none of the 12 structurally incomplete Experiment 6 executions. The public completeness invariant blocked each confident ruling and made escalation effective in all 12. The same enforcement path later stayed silent on every complete pilot execution. This rules out the degenerate implementation that escalates every packet, but only within the declared contract family."

Study 5 (Composition): C7 composed dependency gating, receipt rework, selective invalidation, and current supervisor closure 3/3; all 60 lifecycle assertions passed. However, "fresh truth, Matrix scoped state, and the fair Phi-compacted control all selected current closure 3/3. C7 therefore supports composition of the assurance lifecycle under a bounded state surface, not a context-accuracy advantage."

Study 6 (Transfer failure — the controlling negative result): "Experiment 7's reviewers agreed on structural completeness for 26/30 synthetic packets (Cohen's κ = 0.659). The frozen contract detected all 23 adjudicated-incomplete packets but blocked all seven reviewer-complete packets: sensitivity 1.00, specificity 0.00, and accuracy 0.767. The validity gate failed. A later diagnostic blocked all 20 reviewer-incomplete packets but admitted only 3/20 reviewer-complete packets, for specificity 0.15 and accuracy 0.575. The paper concludes: The negative result separates mechanism correctness from contract validity. The authored ladders show that Matrix can enforce a declared contract; Experiment 7 shows that the contract's notion of completeness may not transfer to independently produced synthetic packets."

Key Distinctions and Boundaries:

  • Governance is not objective truth: Matrix records accepted institutional state under explicit authority and evidence policies. It does not equate acceptance, consensus, or ledger presence with objective truth.

  • Hard mechanics around soft judgment: Metadata-decidable invariants can be hard, model-independent gates. Applicability, completeness, and conflict adjudication that require open-world interpretation remain soft judgments.

  • The paper explicitly states: These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.

Contributions:

  1. It defines governed execution as a conjunction of decision, execution, and change provenance, and makes each property measurable separately from task correctness.

  2. "It reports a controlling negative result: a deterministically enforced completeness contract achieved perfect sensitivity but zero specificity on the initial role-separated transfer set, showing that mechanism fidelity does not establish contract validity."

  3. "It presents an implemented lifecycle joining versioned authority and facts to decision admission, task issuance, receipt-verified closure, supervisory verdicts, and dependency-scoped recovery, with controlled same-outcome demonstrations across those surfaces."

Scope Limitations: "The principal studies use small, self-authored synthetic fixtures and repeated model seeds. They establish implemented mechanisms, not enterprise prevalence, human-review cost, semantic completeness, or general accuracy gains. The paper does not claim superiority over citation-verifying RAG, objective truth from ledger acceptance, natural detection of every relevant change, or validated process forecasting."

Improvements for AI systems

Improvements to AI Systems:

  1. Separate task correctness from governed execution. AI systems should maintain two distinct evaluation channels: one for outcome accuracy and one for institutional compliance (authority, evidence, change tracking). This prevents a system from being scored as correct when it relies on expired waivers, non-authoritative sources, or unsupported completion claims.

  2. Add a deterministic causal-state layer for provenance. AI systems should include a non-model-dependent component that records decision authority, fact dependencies, and completion evidence. This layer owns stable identity, direct references, transition validity, and dependency traversal—removing these from the model's probabilistic judgment to ensure auditability.

  3. Enforce independent completion evidence. AI systems should reject an agent's own claim of task completion as sufficient proof. Instead, they should require externally checkable receipts (e.g., countersigned documents, validated artifacts) that match the task's declared definition of done, with bounded retry counters for corrections.

  4. Implement selective invalidation on policy changes. When an authority or fact is superseded, AI systems should traverse declared dependencies to invalidate only affected work, rather than re-executing all tasks. This reduces redundant rework (e.g., sixfold reduction in task reexecution) while preserving unaffected proof.

  5. Add structural completeness invariants for refusal. AI systems should include public, metadata-decidable gates that block confident rulings when required inputs (e.g., receipts, approvals) are missing. This enables effective escalation to human supervisors without degenerating into always-escalating behavior.

  6. Distinguish mechanism fidelity from contract validity. AI systems should validate that their enforced contracts (e.g., completeness definitions) actually transfer to real-world inputs. The controlling negative result shows that a perfect-sensitivity contract can have zero specificity on independently produced data, so systems must test contract validity separately from enforcement correctness.

  7. Compose governance primitives into a lifecycle. AI systems should integrate dependency gating, receipt rework, selective invalidation, and supervisory closure into a single workflow, allowing bounded state surfaces for complex institutional tasks (e.g., compliance audits, regulatory filings) with full lifecycle assertions.

  8. Maintain explicit boundaries between hard mechanics and soft judgment. AI systems should use deterministic, model-independent gates for what is metadata-decidable (e.g., receipt presence, dependency existence) while keeping open-world interpretation (e.g., applicability, completeness judgment) as soft, human-reviewable steps—never conflating the two.

What the improved AI system can do:

  • Execute agentic workflows in regulated domains (finance, healthcare, government) where outcomes are correct and provably governed—with auditable decision authority, verifiable completion evidence, and traceable change impact.

  • Automatically invalidate only stale work after policy updates, avoiding costly full re-execution while maintaining compliance.

  • Refuse to self-close tasks without independent evidence, escalating to human supervisors only when structural invariants are violated, not on every packet.

  • Detect when its own completeness contract fails to generalize to new data, prompting contract revision rather than silently producing false approvals.

  • Provide a full provenance trail for every action—what authority governed it, what facts supported it, and what prior work it invalidates—enabling third-party audit without model introspection.

Abstract

Agentic workflows are commonly evaluated by whether they reach the correct outcome. That is insufficient in institutional settings, where a correct action may rely on the wrong authority, an unsupported completion claim, or work made stale by a later change. We define governed execution as work whose decisions, completion, and response to change are supported by inspectable provenance. We present Matrix, a deterministic causal-state layer that records authority and fact dependencies, verifies completion evidence, and selectively invalidates affected work. Across controlled comparisons, governed and direct workflows often reached the same outcomes, but only the governed path consistently preserved governing evidence, refused unsupported closure, and limited recovery to dependent tasks. A role-separated transfer challenge then failed: a deterministically enforced completeness contract severely over-blocked synthetic packets produced outside its authoring context. These results do not establish Matrix as a general accuracy enhancer; they support its primary role as an institutional integrity layer for making agentic work auditable and independently verifiable.

Sources

Related papers