The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "The Acknowledgment Point Is the System".
Elias: An AI audit record is useful only if its durability and trust boundary are explicit, and this research introduces RuntimeGuard-AI V2,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So we're diving into "The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence," which seems to be tackling a real problem around how we can trust AI audit records when durability isn't guaranteed by default. This paper is proposing a system that explicitly separates fast decisions from slow, durable persistence operations so we know exactly what our audit evidence is.
Elias: It sounds like they're building this RuntimeGuard-AI V2 architecture to make the durability part of the interface explicit rather than just an implied feature that might fail during a crash. That distinction between immediate acknowledgment and durable persistence seems central to their approach.
Priya: From a privacy perspective, I'm interested in how they handle the policy source binding, because if we can't trust which policy led to a decision, then all the audit evidence is meaningless. We need strong guarantees that what we recorded actually corresponds to the specific deterministic logic that was running.
Nadia: Exactly! The paper focuses on binding each deterministic policy decision to its exact source bytes and calculating a digest from those bytes, which prevents someone from swapping out the policy mid-stream without changing the record commitment. This is a huge step for accountability.
Elias: And I see them generating a request commitment using SHA-two hundred fifty-six on the request domain combined with an encoding of R, which establishes a protocol-stable reference point for every decision made by the AI system. That makes tracing back very specific.
Priya: But what about the practical side? The paper mentions they use an operating-system-backed exclusive writer lease on the evidence directory to serialize sequence assignments and appends, which suggests they are building in some kind of strong consistency for writes even before we get to the durability modes.
Nadia: Right, that exclusivity mechanism is key because it ensures that a later record can't overtake an earlier one if the append hasn't finished, which prevents those kinds of messy sequence gaps later on. But they still offer these three operational modes: buffered, data sync, and full sync for when you need guaranteed crash survival.
Elias: The distinction between those modes is interesting because it makes durability a machine-checkable part of the interface; you can literally check if `durable` is true or false in the receipt returned, turning it from an implied property into an explicit state.
Priya: I wonder how this affects our ability to perform measurements, since we're looking at data integrity and continuity; if there are gaps or duplication issues, how does that impact the overall data quality?
Nadia: The paper addresses that by having a state machine for commits where it evaluates the policy, allocates a global sequence number, constructs the record commitment with its checksummed frame, and then applies either none, data sync, or full sync before signing. If anything goes wrong during append or synchronization, the engine enters a fail-stopped state to prevent those unrecoverable gaps.
Title and authors: Elias: That strict ordering enforced by the state machine is what guarantees that if you replay an identifier and commitment exactly, you get the original decision; reusing that same identifier with different content gets rejected.
Priya: And then they move into this separate attestation path, creating chained Merkle trees from a non-empty, single-policy record range to build a root hash that binds the entire historical chain together for external auditors. That sounds like a way to verify the integrity of an observed epoch chain.
Nadia: It's powerful because it allows an auditor to verify things asynchronously by checking this Merkle root against an externally obtained key, resolving the record commitment and verifying policy equality across the whole range. It proves integrity of that epoch chain without needing access to every single individual record.
Elias: The security model seems thorough too, considering they explicitly consider interruptions during append, corruption of complete frames, replayed requests, conflicting reuse of a request identifier, policy or record tampering, invalid Merkle proofs and even attacker-generated signatures under an untrusted key.
Priya: That level of consideration for tampering is important when we think about the broader impact; if this system can reliably prove which AI decision was made and under what conditions, it opens up new avenues for safety analysis in complex robotic systems.
Nadia: Indeed, and the performance trade-offs are quite concrete; they measured a durability-latency trade-off, showing that policy evaluation is fast at sub-microsecond speeds.
Elias: But when you move to constructing, appending, and signing buffered evidence, the median latency jumps to about one hundred forty-one point eight seven five microseconds under specific testing conditions on an Apple M4 Pro with four worker threads and two thousand forty-eight-byte prompts <ref:2608.17176#pg0,on an Apple M4 Pro>.
Priya: That latency jump is significant; we need to know if that trade-off is acceptable for real-time decision making versus when we need the guaranteed durability of the data sync or full sync modes which take much longer, around sixteen seconds in some cases.
Nadia: The paper shows that synchronized throughput remains near two hundred forty-three requests per second, but synchronized median latency grows from about four milliseconds at one thread up to sixteen milliseconds at four threads and thirty-two milliseconds at eight threads because sharding doesn't create parallel sequence authorities.
Elias: That observation that sharding distributes files but doesn't create parallel authorities is a very practical insight for anyone building on this, as it tells you where the bottlenecks are in terms of sequencing control.
Priya: And we can also see how the cost of generating proofs scales; constructing and signing an epoch with one hundred thousand records takes about ninety-seven milliseconds, which shows that the seal time grows substantially with batch size.
Title and authors: Nadia: The paper also points out that recovery is near-linear in retained records, taking about six hundred sixty-five milliseconds to open and validate one hundred thousand records combined with a total measured recovery path of about nine hundred sixty one milliseconds.
Elias: Looking at the limitations mentioned, the authors are clear that they're testing on a "one deterministic regex policy fixture," and they also don't measure network service latency or multi-host availability, which means these findings aren't directly applicable to production SLOs across a distributed infrastructure.
Priya: And another crucial limitation is that the system doesn't authenticate caller-supplied model identity or user identity; it only records commitments to those assertions, and they also state plainly that checksums alone don't resist a privileged operator who could rewrite complete records.
Nadia: So, while this framework provides explicit durability receipts and verifiable epoch chains for audit evidence, we still need external witnesses for fork detection or trusted execution environments to fully secure the underlying execution integrity of the AI itself.
Elias: That’s the gap they've identified; no cryptographic relation proves policy or model execution on its own, so stronger guarantees require those external layers you mentioned.
Priya: It seems like this paper provides a very solid foundation for building verifiable audit trails, and even with these limitations, it gives us a clear blueprint for how to design systems that can produce durable receipts for AI decisions.
Nadia: We've covered the title and authors of "The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence," walked through the core summary of what RuntimeGuard-AI V2 actually does, discussed their proposed improvements to make durability explicit, and looked at their concluding thoughts on performance versus practical limitations.
Elias: It's clear that this research is focusing on turning the abstract concept of audit evidence into a machine-checkable state by rigorously defining synchronization boundaries and commitment protocols.
Priya: For us, the implications suggest a path toward more transparent AI systems where we can definitively prove what was decided and under which constraints, which is essential for building trust in complex applications.
Nadia: It’s exciting to see this level of detail on how to manage the trade-off between low latency and guaranteed crash survival when dealing with AI audit logs.
Elias: We've seen how the commitment state machine enforces strict ordering, and the separate attestation path for Merkle epochs is a clever way to handle verifying historical integrity asynchronously.
Priya: I think what sticks is the explicit interface for durability—buffered versus data sync versus full sync—which lets operators choose their audit path based on their immediate needs.
Nadia: So, if we look ahead, this paper gives us a lot of material to work with when designing the next generation of verifiable AI systems, focusing on making trust boundaries totally transparent.
The paper's summary: Nadia: So, to recap, this paper is proposing a system that makes it explicit whether an AI decision has been recorded durably or just buffered for quick acknowledgment, which is a big deal for audit trails.
Elias: Yeah, and what caught my eye was how they turn the durability bit into a machine-checkable state in the interface itself rather than leaving it as some hidden property.
Priya: From my angle, this really changes how we think about verifying historical AI behavior; instead of just looking at a log file, we can actually check if that specific decision was recorded under what conditions.
Nadia: Exactly! It moves us from just having a record to actually having a verifiable receipt that tells us the system's commitment level at that moment.
Elias: And the way they tie the policy source bytes directly into the commitment digest is something I find interesting because it anchors the decision to a specific piece of logic.
Priya: That specificity is what makes it useful for privacy checks; we can trace a specific output back to a precise rule set, which helps us understand potential bias or misuse.
Nadia: It’s powerful because it prevents someone from later claiming they followed policy A when they actually ran policy B, because the source bytes are hashed in there.
Elias: And the protocol for creating that request commitment using SHA-two hundred fifty-six on the request domain and the encoded request is a solid cryptographic foundation for tracing.
Priya: The fact that they separate this into a synchronous commit path and an asynchronous attestation path means we get both fast feedback and deep, verifiable historical proof at different times.
Nadia: That separation is key; it lets us use the system quickly for real-time needs while still having the heavy lifting for long-term audit trails happening in the background.
Elias: I agree, that layered approach to verification is smart; it doesn't force you to wait for a full sync just to check if a record even exists.
Priya: So, when we look at the results, they show that durability really matters; policy evaluation itself is quick, but getting that durable receipt involves a noticeable latency increase depending on how much data you need to sync.
Nadia: That latency trade-off is something I'm thinking about; if you need an instant decision every time, are we willing to accept that slower path for auditability?
Elias: The paper provides some concrete numbers on that performance gap, showing how much it takes to move from a simple acknowledgment to a fully synchronized durable state.
Priya: The implication for me is that this system offers a way to measure the cost of different levels of trust; we can quantify exactly how much time and effort goes into achieving higher assurance.
Nadia: It gives us the metrics we need to make those tough operational decisions about where we prioritize speed versus absolute certainty in our AI deployments.
Elias: And that ties back to my point on the cryptographic assumptions—the proof relies heavily on the integrity of that initial policy source and the sequencing mechanism being perfectly enforced.
Priya: So, essentially, this work gives us a way to build verifiable history by making the durability cost transparent and tying every decision back to its exact origin.
Nadia: It sounds like a lot of practical tools for building better AI accountability systems across the industry.
Elias: Indeed; the next step is figuring out how to secure that entire chain against an adversary who might try to tamper with the sequence assignments themselves.
Priya: Next time, we should really dig into those limitations they mentioned regarding external witnesses and trusted execution environments, because that’s where the real security challenge lies.
The paper's improvements: Nadia: So, we’re looking at how the authors suggest they can make this system even more robust by adding several layers of improvement to RuntimeGuard-AI V2, which is exciting stuff.
Elias: Yeah, I was particularly interested in their idea to generate Ed25519-signed receipts for every single committed record; that seems like it would provide a very strong cryptographic binding.
Priya: And the shift toward chaining those records into Merkle epochs, where each epoch root hash summarizes a whole block of historical decisions, sounds like it really solidifies the integrity check.
Nadia: It’s about creating this end-to-end verification path so an auditor doesn't have to trust any single piece of evidence but can verify the entire chain structure.
Elias: That chaining mechanism is strong because it links the policy descriptor and sequence range into that root hash, making it hard to forge a historical epoch without knowing all the preceding hashes.
Priya: From a privacy standpoint, having these verifiable epoch statements allows us to confirm that a specific AI behavior wasn't accidentally included in an unauthorized update or modification of the system's policy over time.
Nadia: Exactly! It moves us from checking one record to proving the entire contiguous history is sound, which is exactly what we need for compliance and safety investigations.
Elias: Also, they propose a strict restart validation logic where the engine checks for sequence continuity and verifies all framed records before it lets the system resume operation.
Priya: That restart logic addresses data corruption directly; it means if something gets damaged during a write operation, the system won't just silently ignore it but will reject the bad state.
Nadia: I’m also interested in their suggestion for a cost-aware mechanism that reports the latency associated with each durability mode, like buffered versus full sync.
Elias: That is practical engineering; knowing exactly how much time and resources you use to get a certain level of audit certainty helps you choose the right path for your specific application needs.
Priya: It allows us to measure the trade-off between real-time performance and deep compliance assurance, which is vital when deploying these complex AI systems in sensitive environments.
Nadia: So they’re not just building a tool; they’re giving us a clear framework for making informed decisions about how much trust we need to place in our audit trails at any given moment.
Elias: And that points toward the future, because securing those external witnesses and trusted execution environments is what will finally give us the kind of proof needed to verify the integrity of the AI itself.
Priya: That’s where our next big question should be—how do we actually implement those external witnesses reliably so they don't become a single point of failure?
Nadia: It’s a tough challenge, but I think this paper sets up the necessary structure for us to start asking those harder questions about securing the underlying execution integrity.
Conclusion: Nadia: So, to wrap things up, this paper on "The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence" shows us a concrete way to formalize durability in AI audit logs by making synchronization explicit and binding policy decisions to source bytes.
Elias: I think what really stands out is how they’ve constructed that commitment state machine with its strict sequencing and fail-stopped states, which provides a solid cryptographic anchor for the evidence.
Priya: From a measurement standpoint, the data clearly shows that while there's a performance hit for durability, we can now quantify exactly what level of historical certainty we're paying for when auditing AI output.
Nadia: It means that when we deploy AI systems, we have a quantifiable way to choose between fast logging and guaranteed persistence based on our operational needs.
Elias: And the idea of the separate attestation path generating Merkle epochs for external verification is a smart way to handle historical integrity without needing every single record in one place.
Priya: It opens up a new avenue for measuring compliance; we can finally produce metrics that prove not just what the AI output was, but exactly how it was decided and stored.
Nadia: It’s powerful because it gives us a verifiable receipt that ties the decision directly to the policy, which is something we desperately need in regulated industries.
Elias: We’ve seen how they handle assumptions regarding policy integrity and sequencing, and while they flag limitations on external witnesses, that sets a clear direction for where future cryptographic hardening needs to go.
Priya: I just want to emphasize that the practical trade-offs discussed in the performance section really matter; we can't forget those latency figures when designing systems for real-time applications.
Nadia: Exactly; knowing those costs lets us design better systems, whether we’re aiming for high-throughput logging or critical compliance recording.
Elias: Overall, this work on "The Acknowledgment Point Is the System: Durable Policy-Decision Receipts for AI Audit Evidence" is a solid step toward making AI decision provenance a standard feature rather than an optional add-on.
Priya: It gives us tools to move beyond simply observing what an AI does and start rigorously measuring and proving its behavior over time.
Nadia: I think the real impact here is moving the conversation from 'Can we trust this record?' to 'What level of trust do we need, and what is the verifiable cost of achieving it?'
Elias: And that's exactly what we need to figure out next, focusing on those external witnesses and operational key lifecycle issues they mentioned in their conclusion.
cs.CR, cs.AI
Submitted: 2026-08-17
Updated: 2026-10-05
Comments: 6 pages, 4 figures, 2 tables. Code and release artifacts: https://github.com/neerazz/RuntimeGuard-AI/releases/tag/v2.0.0
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: An AI audit record is useful only if its durability and trust boundary are explicit, and this research introduces RuntimeGuard-AI V2, a system that binds deterministic policy decisions to explicit
Key concepts
- Durability Modes
- The system offers three operational modes: 'buffered' (non-durable), 'data sync' (requires data synchronization before being durable), and 'full sync' (synchronizes all data before being durable). This makes the durability status an explicit, machine-checkable part of the interface rather than an implied property.
- Request Commitment
- A commitment is a cryptographic hash generated from the request domain and encrypted content. It uses fixed-width lengths and a fixed field order to ensure consistency. This commitment serves as a unique identifier for a specific policy decision, allowing for exact replay verification.
- Commit State Machine
- This machine enforces strict ordering of all writes. It manages sequence allocation, shard selection, record construction, and synchronization operations in a defined state flow. If any step fails during append or synchronization, the engine stops to prevent creating unrecoverable gaps in the evidence chain.
Terminology
Summary
An AI audit record is useful only if its durability and trust boundary are explicit, and this research introduces RuntimeGuard-AI V2, a system that binds deterministic policy decisions to explicit synchronization boundaries to create durable receipts for AI audit evidence.
How it works
The core mechanism revolves around the distinction between immediate acknowledgment and durable persistence. The system exposes three operational modes: buffered,
which returns a receipt with durable=false
; data sync,
which requires a call to sync data before returning durable=true
; and full sync,
which calls sync all before returning durable=true.
This makes durability a machine-checkable part of the interface, turning it from an implied property into an explicit state.
The protocol involves several steps:
-
A compiled policy contains
canonical source bytes
and a descriptor (id, version, digest), where the digest is calculated asSHA−256(source).
-
A request commitment is generated as
HR = SHA− 256(request-domain∥ enc(R)), where enc uses fixed-width lengths and a fixed field order.
-
The engine maintains an
operating-system-backed exclusive writer lease on the evidence directory,
with a mutex serializing sequence assignment and append to ensurea later record cannot overtake a sequence whose append has not completed.
Commit State Machine
The state machine dictates strict ordering for writes. For every new request, the engine evaluates the policy, allocates the next global sequence, selects a shard sequence mod K, constructs a protocol-stable record commitment, appends one checksummed frame, and applies the selected synchronization operation before signing the receipt. If append or synchronization fails,
the engine enters a fail-stopped state.
This prevents creating gaps that later recovery must hide or reject. Exact replay of an identifier and commitment returns the original decision; reuse of the identifier with different committed content is rejected.
Signed Epochs and Audit
The attestation path operates separately from the synchronous commit path, focusing on asynchronous batch attestation. It sorts a non-empty, single-policy, contiguous record range
to build a SHA-256 Merkle tree.
An epoch statement binds the root hash, logical count, sequence range (first and last), policy descriptor, predecessor statement hash, and signer key. An auditor verifies this by checking the signature against an externally obtained key,
resolving the record commitment, verifying the epoch signature and policy equality, deriving proof indices from the sequence range, and verifying Merkle inclusion. This proves integrity of an observed epoch chain.
Performance Trade-offs
The research measures a durability–latency trade-off,
not a free asynchronous path. Under specific conditions (Apple M4 Pro at four worker threads and 2,048-byte prompts), the performance metrics are:
>(RQ1) durability dominates the cost.
Policy evaluation alone is sub-microsecond,
but Constructing, appending, and signing buffered evidence raises median latency to 141.875 µs.
Data and full synchronization raise median latency to 16,018.646
and 16,009.136 µs,
respectively.
>(RQ2) serialization saturates rather than scales.
Per-record synchronized throughput remains near 243 requests/s across worker counts.
However, synchronized median latency grows from approximately 4 ms at one thread to 16 ms at four and 32 ms at eight,
indicating that Sharding distributes files but does not create parallel sequence authorities.
>(RQ3) epoch construction scales with batch size.
At 100,000 records, building and signing an epoch takes 96,954.958 µs (97.0 ms),
showing that Seal time grows from 0.116 ms at 100 records to 97.0 ms at 100,000 records.
>(RQ4) recovery is near-linear in retained records.
Opening and validating 1,00,۰۰۰ records takes 665.483 ms,
with the combined measured recovery path being 961.0 ms.
Limitations
The prototype evaluates a one deterministic regex policy fixture,
and it does not measure network service latency, multi-host availability, or production SLOs. Furthermore, the system does not authenticate caller-supplied model identity or user identity; it only records commitments to those assertions. Checksums alone do not resist a privileged operator who rewrites complete records. The paper explicitly states that no cryptographic relation proves policy or model execution.
Stronger guarantees require mechanisms such as external witnesses for fork detection, trusted execution or proof systems for evaluator integrity, and operational key lifecycle for long-lived trust.
Improvements for AI systems
Here are the specific improvements an AI system could make by implementing the concepts from RuntimeGuard-AI V2, and what those improvements would allow it to achieve:
-
The AI system can implement a
durable policy-decision receipt
mechanism that explicitly separates fast, non-durable acknowledgments from slow, durable persistence operations. This allows the system to choose between low latency (for real-time decisions) and guaranteed auditability (when durability is required), making the trade-off explicit in its interface rather than hidden in a crash. -
The system can enforce strict source-bound policy binding by requiring that every compiled policy contains canonical source bytes and a deterministic digest derived from those bytes. This prevents callers from attaching arbitrary logic to unrelated policy decisions, ensuring that the executed decision is provably tied to the specific, audited policy artifact used at runtime.
-
The system can utilize a single-writer commit protocol with explicit synchronization semantics (via modes like 'data sync' or 'full sync'). This ensures that evidence is only acknowledged as durable after a verifiable host synchronization call succeeds, guaranteeing that the recorded state survives immediate process or power failures during the critical write phase.
-
The system can generate Ed25519-signed receipts for every committed record, binding the request identifier, commitment, sequence number, durability bit (true/false), and verification key directly to the decision. This provides an immutable audit trail proving exactly when a decision was acknowledged and under which durability constraints.
-
The system can construct chained Merkle epochs from committed records, where each epoch is sealed with a root hash that includes the policy descriptor, sequence range, and predecessor statement hash. This allows for end-to-end verification by an auditor to confirm the integrity of a contiguous block of historical decisions without needing to trust any single intermediary.
-
The system can be designed with strict restart validation logic upon reboot. When restarting, the engine validates framed records (including magic bytes and checksums), ensures sequence continuity (checking for gaps or duplicates), and verifies that all committed records adhere to the expected format and policy digest before resuming operation, rejecting corrupted or incomplete states rather than silently accepting them.
-
The system can provide a cost-aware mechanism for auditing by reporting the specific latency and throughput costs associated with each durability mode (policy only, buffered evidence, data sync, full sync) based on experimental benchmarks. This allows operators to make informed decisions about which audit path to use for different operational needs (e.g., using the fast path for high-volume logging versus the slow path for critical compliance logs).
-
The system can be enhanced with a separate
attestation path
that groups committed records into signed epochs, allowing external auditors to verify inclusion proofs against an externally obtained key, proving that specific historical decisions were included in a valid, chained history without requiring full access to the internal state of every record.
Abstract
An AI audit record is useful only if its durability and trust boundary are explicit. Returning a guarded decision before any durable write minimizes latency, but it cannot guarantee that evidence survives an immediate crash. We rebuild RuntimeGuard-AI around this constraint. The resulting research prototype binds each deterministic policy decision to the exact policy source, commits a privacy-minimizing record at a caller-selected synchronization boundary, and returns an Ed25519-signed receipt that states whether that boundary completed. After restart, the engine validates framed records, manifests, shard placement, sequence continuity, and replay identity. A separate attestation path groups committed records into chained, signed Merkle epochs that an auditor verifies with an externally obtained key. On an Apple M4 Pro at four worker threads and 2,048-byte prompts, buffered signed evidence reaches 27,193 requests/s with 141.9 microseconds median latency. Per-record data and full synchronization reduce throughput to approximately 242 requests/s and raise median latency to 16.0 ms. Sealing a 100,000-record signed epoch takes 97.0 ms. The result is a measured durability-latency trade-off, not a "free" asynchronous audit path. The prototype does not prove model execution, prevent a compromised signer from forking history, or establish legal conformity.
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs