Beyond Memory: A Transactional Continuity Kernel for Long-Lived AI Agents
Jun He, Deying Yu
OpenKedge.io
cs.MA, cs.AI
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 9 pages, 6-page appendix, 15 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
Terminology
Summary
arXiv: 2608.11632v1 [cs.MA] 12 Aug 2026
The paper addresses a fundamental issue in persistent AI agents: "Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an explicit control plane, unmediated updates by models, tools, and background workers risk stale overwrites, un-audited exposures, and self-authorizing privilege escalation."
The authors argue that agent state governance is an infrastructural activation problem, defining continuity as an unbroken, authorized lineage of accepted branch heads.
They define continuity as "an unbroken, authorized lineage of accepted heads on one branch. This is infrastructural continuity—a guarantee about state activation and lineal provenance—not a philosophical claim about agent consciousness or behavioral identity."
The core thesis: agent state governance is an infrastructural activation contract problem.
The paper presents the Continuity Kernel (CK), an activation contract that decouples off-commit candidate evaluation from atomic state activation.
CK establishes a strict boundary between candidate object preparation and authoritative state activation.
Agent state is partitioned into isolated branches identified by (sid, bid). The storage layer maintains multiple functional namespaces (Accepted, Candidate, Attempt, Quarantine, Projection), but storage presence alone does not confer authority: an object is authoritative if and only if it is reachable from the current branch head committed by CK.
The branch directory maintains the complete-head type h and branch-row type B[k]:
h = ⟨sid, bid, seq, root, v, parentRef, lineageRef⟩
B[k] = ⟨h, status, writer, epoch, dirSeq, handoff, createdFrom⟩
Models, tools, memory services, and operators are untrusted proposers. Evaluators and approvers issue signed evidence but cannot advance a head.
A signed proposal is summarized by τ = ⟨pid, target, expected, ops, a, req⟩.
Critically, authorization evaluates against the pre-state context A
— a proposal cannot authorize itself. The candidate seal binds the proposal, evidence, interpreter, and component roots. Activation recomputes these bindings against the serialized branch state.
The safety claims depend on nine explicit assumptions (A1–A9): (1) Boundary & Cryptography (A1–A3); (2) Atomic Serialization (A4); (3) Complete Context & Time (A5–A6); (4) Lifecycle Order & Non-Recycled Scopes (A7–A8); and (5) Verification Objects (A9).
Under those assumptions, the contract has four consequences:
-
Exact succession — one complete predecessor has at most one accepted successor
-
Pre-state admission — authorization and freshness checked against serialized predecessor, never proposed context
-
Stable execution identity — one terminal disposition per proposal identifier, at most one accepted binding per effect identifier
-
Lifecycle isolation — branch creation, fencing, handoff, migration, restoration preserve branch scope and current authority
The protocol separates slow, fallible preparation from one short serialized decision.
The activation transaction evaluates an ordered vector of 17 predicates (ActCheck): TargetLookup, Allocation, ExpectedHead, Lifecycle, HandoffRevision, AcquisitionAuth, RequirementCoverage, VersionBinding, WitnessBinding, CandidateBinding, EffectBinding, FreshnessInitial, Authorization, Decision, AuthorityTransition, FreshnessFinal, WellFormedness.
The kernel evaluates this vector strictly in the Appendix A order and accepts only an all-Pass vector.
Each stage has exactly one predicate; no aggregate guard can move a failure across stages.
For an authenticated, owner-bound identifier, the kernel records exactly one of four terminal dispositions:
-
Commit — install the exact successor and all acceptance metadata atomically
-
Reject — permanent validation or policy failure; retain no successor
-
Quarantine — retain sealed material outside the authoritative head
-
Defer — trusted acquisition reports required evidence unavailable
Every disposition is terminal for its proposal identifier. Resubmission after Defer or Quarantine uses a new identifier linked to the old attempt.
The construction follows a strictly cycle-free order:
Pb → WellFormedness(Pb) → rC → h′ → Uaccept
The complete accepted unit is:
Uaccept = X′, Γ′, h′, B′[k], rC, O[pid]
Receipt verification has four increasing evidence levels:
-
Structural — canonical syntax, typed bindings, valid signature
-
Attested — trusted kernel key attests first terminal stage and reason
-
Replay — retained inputs and pinned versions reproduce the stated decision
-
Inclusion — certified snapshot or log proof places receipt and outcome in durable state
The distinction matters because a kernel may sign a receipt before its storage transaction aborts.
A new branch has no predecessor. Its genesis record binds the authenticated creation request, dξ, fresh identifier, initial roots/writer, and absence proof.
Authorization uses serialized creation context ξk, never nonexistent predecessor state.
A handoff uses a monotonically versioned directory record.
The state machine has states: Prepared, SourceFenced, Aborted, TargetActive. Actions include Prepare, Abort, Fence, Retarget, Activate. The gap between Fence and Activate is intentional. After fencing, no writer is authorized; therefore a crash cannot expose two active writers.
Migration names source and target schemas plus a pinned deterministic migrator.
Restoration creates a new successor: X′ = RestoreM(Xc, Mµ(Xk)) with Γ′ = Γc. Paths in M take checkpoint values after migration; all other paths take current values. Authority, revocation, writer epoch, and lifecycle fields are protected and therefore remain current.
The paper evaluates CK through four research questions (RQ1–RQ4). The authors construct an executable bounded model in Python (artifacts/bounded model.py)
performing exhaustive breadth-first search (BFS) state exploration over a finite abstraction of the kernel specification.
The model state space includes: one active branch directory and state store, shared admission and creation contexts, honest and adversarial principals, 13 proposal IDs, 4 effect IDs, source and target writers, and integer versioning.
Key results:
-
2,808,230 reachable states (depth 7)
-
5,526,474 unique state-changing transitions
-
8,880,248 total transition attempts (including idempotency checks and retries)
-
Zero invariant violations across all reachable states
-
100% of named protocol coverage witnesses reached
Transition outcome distribution (of 8.88M attempts):
-
IdReuseConflict: 18.84%
-
Commit (Head Advance): 10.33%
-
ReclaimedEffects: 9.48%
-
RejectStaleHead: 8.68%
-
PrefixNotFinal: 8.24%
-
Defer (Missing Evidence): 4.71%
-
RejectUnauthorized: 4.46%
-
Other dispositions: 13.73%
Performance: kernel transition evaluation achieves a throughput of 31,132 transitions/sec with an average evaluation latency of 32.12 µs per transition.
-
Pre-State Authorization Safety (RQ1):
For every committed transition, authorization is checked against the predecessor authority context Γ (or genesis context ξk for BranchCreate).
Self-authorizing proposals were rejected 395,905 times (4.46% of transitions) with RejectUnauthorized. -
Execution Identity & Replay Safety (RQ2):
Each proposal identifier pid has at most one durable disposition O[pid], and each effect identifier eid is bound at most once in E[eid].
-
Writer Fencing & Lifecycle Isolation (RQ3):
Once the source writer is fenced (SourceFenced), any subsequent write attempt by the old writer w0 fails with StaleRevision or WriterEpoch.
The paper identifies critical gaps between formal model and physical implementations:
-
Abstraction vs. Physical Storage Realization — physical storage crashes, network partitions, storage-engine bugs fall outside the model's abstract state space
-
Signature vs. Durable Inclusion — signed receipts prove cryptographic origin but not transactional durability
-
External Side-Effect Atomicity — CK cannot make irreversible remote side effects atomic with internal state updates
-
Performance & Latency Trade-offs — 17 sequential stage checks, double freshness validation, prospective unit construction add overhead
Three primary safety properties are proven (conditional on A1–A9):
Proposition 1 (Owner-bound stable outcome): Under A1–A4 and A8, only the principal named by an active allocation may create O[pid], and at most one disposition becomes durable for that identifier.
Proposition 2 (Single exact continuation): Under A1–A4, at most one proposal commits from one complete predecessor or serialized target absence, and the installed state is its sealed candidate.
Proposition 3 (At-most-once accepted effect): Under A1, A4, and A8, each effect identifier has at most one accepted binding, including after online effect records are reclaimed.
The paper notes: The propositions do not claim semantic correctness: a policy may approve a false memory, and an idempotent intent may still cause a non-idempotent remote action if its connector is faulty.
The paper positions CK relative to:
-
Agent memory systems: Generative Agents, MemGPT, LoCoMo, A-MEM, Mem0, MemOS — these
motivate governed memory
-
MemTX (closest agent-memory transaction system):
stages evidence-bearing beliefs under a snapshot, gates actions, and propagates repairs after retraction
— CK applies an activation boundary to typed persistent state -
MemTxn:
supplies source-supported memory admission, visible-version selection, snapshots, and journal-based recovery
-
Transactions and retries: optimistic concurrency control, database transactions, Raft consensus, RIFL (Reusable Infrastructure for Linearizability)
-
Effects and evidence: Cordon (task-level semantic transactions), PROV-DM, JSON-LD, RDFC-1.0, Data Integrity
The paper concludes: "For persistent agents, state retention is not authority. Infrastructural continuity requires an authorized lineage of accepted branch heads. CK makes the transition to authority explicit: untrusted components propose typed candidates off-commit; a deterministic control plane validates an exact predecessor head and pre-state authority (or typed absence and genesis context); a short activation transaction records one stable disposition; and only Commit advances the branch head."
The model validates logical protocol consistency within its finite bounds; it does not establish unbounded mathematical correctness, physical storage driver conformance, or throughput performance under hardware faults.
"The core contribution of this work is the activation contract itself—providing a principled systems foundation that separates stored retention from authoritative reachability and distinguishes cryptographic receipt signatures from durable transactional inclusion."
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Implement an activation contract layer that separates candidate state preparation from authoritative state activation. The AI system maintains a branch directory with complete-head types and branch rows, ensuring only reachable states from committed heads are treated as authoritative.
-
Add a 17-predicate validation vector (ActCheck) that evaluates proposals in strict order before any state transition. This includes target lookup, allocation, expected head verification, lifecycle checks, handoff revision, acquisition authorization, requirement coverage, version/witness/candidate/effect binding, freshness validation (both initial and final), authorization, decision, and authority transition checks.
-
Enforce pre-state authorization where all permission checks evaluate against the serialized predecessor state context, never the proposed context. This prevents self-authorizing privilege escalation.
-
Implement four terminal dispositions (Commit, Reject, Quarantine, Defer) with strict at-most-once semantics per proposal identifier. Resubmission requires new identifiers linked to old attempts.
-
Add writer fencing state machine with states: Prepared, SourceFenced, Aborted, TargetActive. The intentional gap between Fence and Activate prevents crash-induced dual-writer scenarios.
-
Implement effect binding at-most-once — each effect identifier binds exactly once, even after online effect records are reclaimed.
-
Add four-level receipt verification: Structural (syntax/signature), Attested (kernel key attestation), Replay (input reproduction), and Inclusion (durable log proof).
-
Use typed absence proofs for genesis records — new branches bind authenticated creation requests with absence proofs rather than nonexistent predecessor state.
-
Maintain long-lived agent state across months/years without stale overwrites or unauthorized state mutations, even when multiple models, tools, and background workers propose concurrent changes.
-
Guarantee exactly one accepted successor per predecessor head — the system can prove that no two conflicting state versions were ever both committed from the same branch point.
-
Provide auditable, replayable decision receipts — every state change has a cryptographic receipt verifiable at four evidence levels, allowing forensic reconstruction of any historical state.
-
Recover safely from crashes during writer handoffs — the fencing protocol ensures no window exists where two writers could both believe they are active.
-
Reject self-authorizing proposals — the system blocks any proposal that attempts to grant itself permissions or modify its own authorization context.
-
Handle missing evidence gracefully — when required evidence is unavailable, the system defers rather than guessing, preserving safety while allowing retry with new identifiers.
-
Maintain stable execution identity — each proposal and effect has exactly one terminal disposition, preventing replay attacks or double-execution of side effects.
-
Support migration and restoration while protecting authority fields — schema migrations preserve current authority, revocation, and writer epoch even when restoring from checkpoints.
-
Provide performance guarantees — with 31,132 transitions/sec throughput and 32.12 µs average latency, the system can handle high-frequency agent state updates without becoming a bottleneck.
-
Prove safety properties — the system can demonstrate (under stated assumptions) that it satisfies exact succession, pre-state admission, stable execution identity, and lifecycle isolation invariants across all reachable states.
Abstract
Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state. Without an explicit control plane, unmediated updates by models, tools, and background workers risk stale overwrites, un-audited exposures, and self-authorizing privilege escalation. We argue that agent state governance is an infrastructural activation problem, defining continuity as an unbroken, authorized lineage of accepted branch heads. We present the Continuity Kernel (CK), an activation contract that decouples off-commit candidate evaluation from atomic state activation. Untrusted components propose typed changes against an exact predecessor head or typed absence. A short activation transaction revalidates ownership, pre-state authority, freshness, and effect uniqueness, recording one stable disposition (Commit, Reject, Quarantine, or Defer). Only Commit atomically advances the branch head and installs the complete accepted unit (state, authority, lineage, effects, outcome, and receipt). A bounded executable model verifies the protocol across 2,808,230 reachable states and 5,526,474 state-changing transitions with zero invariant violations.
Sources
- MemGPT: Towards LLMs as Operating Systems
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- MemOS: A Memory OS for AI System
- MemTX: Transactional Belief Commit for Stateful Agent Memory
- MemTxn: A Transaction Boundary for Source-Supported Updates and Complete-State Recovery in Agent Memory
- Temporary Authority, Permanent Effects: Commit-Time Authorization for LLM Agents
- Cordon: Semantic Transactions for Tool-Using LLM Agents
Related papers
- Highway Congestion Reduction through Reinforcement Learning Based Eulerian Headway Control
- You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents
- Deny Without Disabling: Authorization-Paired Evaluation and Control for Multi-Agent Systems
- MA-SAPO: Multi-Agent Reasoning for Score-Aware Prompt Optimization
- PeroMAS: A Multi-agent System of Perovskite Material Discovery
- StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning