Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents
Qingdao Guodongxiansheng Network Technology Company Limited
cs.AI
Submitted: 2026-08-12
Updated: 2026-08-26
Comments: 23 pages, 2 figures, 11 tables. Includes a sealed end-to-end governed-memory service evaluation with an ungoverned local Qwen2.5-7B comparison
Code: https://github.com/kzkz137806/aethmere-os
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 95/100
The gist: Governed Persistent Memory (GPM) is introduced as an auditable bitemporal state-transition model for long-horizon agents, addressing the problem that retrieval alone does not determine whether
Terminology
Summary
Governed Persistent Memory (GPM) is introduced as an auditable bitemporal state-transition model for long-horizon agents, addressing the problem that retrieval alone does not determine whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. The model includes source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head.
On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies (flat conflict-preserving) matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches.
These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.
Improvements for AI systems
Improvements to AI Systems:
-
Add a bitemporal state-transition layer to long-hhorizon agents, so every retrieved fact carries a lifecycle state (active, superseded, retracted, deleted) and a validity interval. The improved system can refuse to generate claims from stale or revoked records, and can explicitly state
this was true until [date], then retracted
instead of presenting outdated information as current. -
Implement source-bound admission control for all ingested knowledge. The system only accepts facts with a verifiable origin and a cryptographic hash, preventing unverifiable or conflicting data from entering the agent's memory. The improved system can trace every claim back to its exact source and timestamp, enabling auditability and dispute resolution.
-
Add fail-closed structured release for answer generation. When the agent produces a response, it must first construct a fresh, consistent view of its memory at a single verified head, then check that the claim is supported by active, non-conflicting records. If any supporting record is missing, superseded, or revoked, the system outputs
cannot confirm
rather than guessing. The improved system can guarantee that every answer is either fully supported by current evidence or explicitly flagged as unverifiable. -
Introduce conflict isolation and non-revival rules in the memory store. Once a fact is retracted or deleted, no future retrieval or inference can resurrect it, and contradictory records are quarantined so they cannot jointly support a claim. The improved system can maintain logical consistency over time, even when new information contradicts old data, and will never produce an answer that relies on a fact that was officially withdrawn.
-
Add exact claim closure verification before any output. The system checks that the set of supporting records for a claim is complete and that no other active record contradicts it, using a deterministic, hash-frozen contract. The improved system can certify that its answer is not only plausible but is the unique logical consequence of the current, verified state of its knowledge base—eliminating hallucination from incomplete or mixed-version retrieval.
-
Use a governed lane for high-stakes queries (e.g., medical, legal, financial). The system runs a separate, auditable inference path that applies the above rules, while a faster ungoverned path handles casual queries. The improved system can automatically route critical questions to the governed lane, ensuring that any answer with real-world consequences is fully traceable and conflict-free, while maintaining speed for low-risk interactions.
What the improved AI system can do: It can answer questions with provable consistency over time, never silently contradicting itself or relying on revoked data, and can produce an audit trail for every claim. It can handle long-horizon tasks (e.g., multi-day research, ongoing project management) where facts change, by automatically updating its internal state and clearly marking what changed and when. It can also generate certified
answers that are guaranteed to be free of internal contradictions and based only on the latest, verified evidence—useful for compliance, journalism, and decision support.
Abstract
Long-term agent memory is usually treated as select--store--retrieve, but retrieval does not decide whether contradictory, superseded, retracted, deleted, or stale records may support an outgoing claim. We introduce Governed Persistent Memory (GPM), an auditable bitemporal state-transition model with source-bound admission, derived lifecycle state, current public barriers, and fail-closed structured release. Five executable clauses cover ledger integrity, source binding, conflict isolation, non-revival after retraction or deletion, and exact claim closure over a fresh view at one verified head. On a prespecified hash-frozen 3,600-case GPM-ReleaseBench, GPM matches all complete outcomes; the strongest of three intentionally simple complete policies matches 1,800/3,600 and makes unmatched releases on 50% of violation cases. A separate sealed end-to-end service evaluation exercises real ingestion and release across eight query families. In its publicly disclosed V3 arm, the governed lane is correct on 2,400/2,400 clusters versus 600/2,400 for ungoverned local Qwen2.5-7B; it repairs all 1,800 baseline failures with no regression (one-sided 95% lower bounds 99.875% and 99.834%). A later V5 reseal over Chinese- and English-command arms, with generation-date pinning and no post-freeze reducer amendment, again obtains 2,400/2,400 per arm. A production-code-independent finite model explores 331,776 semantic and 1,990,656 query states without a full-contract counterexample, and a 100,000-trace three-engine differential yields zero mismatches. These are bounded contract and implementation results, not open-world model accuracy or evidence of world truth. Governed answers in the sealed service evaluation are deterministic service outputs; the 7B result is the ungoverned comparison, not a claim that a language model itself became perfectly accurate.
Sources
- HaluMem: Evaluating Hallucinations in Memory Systems of Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
- Mitigating Provenance-Role Collapse in Long-Term Agents via Typed Memory Representation
- MemOS: A Memory OS for AI System
- MemGPT: Towards LLMs as Operating Systems
- MemConflict: Evaluating Long-Term Memory Systems Under Memory Conflicts
- StateFuse: Deterministic Conflict-Preserving Memory for Multi-Agent Systems
- Augmenting Language Models with Long-Term Memory
- TOKI: A Bitemporal Operator Algebra for Contradiction Resolution in LLM-Agent Persistent Memory
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- MEMOREPAIR: Barrier-First Cascade Repair in Agentic Memory
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection