MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems

arXiv:2605.25002 · cs.CR · Submitted 2026-05-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems".

Jane: The paper was written by the authors from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Summary: Tom: So, we're moving on to talk about the summary of MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems. The authors are explaining that when you lose your logs or your snapshots, traditional methods fail because the same untrusted file holds both the data and any claims about its origin.

Jane: It's like having a diary where you try to prove who wrote it with a signature, but someone else erasing the signature and rewrite all the entries; MemMark gives us a way to verify that without needing external logs.

Lu: The core idea is that instead of putting attribution in mutable metadata, MemMark embeds it into the actual decision-making process of choosing which memory item to update or link.

Meng: This is interesting because it forces the backend LLM—the one doing the thinking—to carry a secret signal within its native function, so we don't have to rely on external tracking systems that might disappear.

Lalam: The summary shows that this method successfully preserves memory utility while embedding this complex, hidden attribution; Lalam feels that it's proving we can have verifiable agents without sacrificing performance.

The Improvements: Tom: Now, let’s look at the improvements and technical specifics of MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems. The key contribution here is that it provides a way to get full payload recovery from just the final memory snapshot, even without any external logs.

Jane: That’s a massive leap because, as we saw in the summary, R3—the snapshot-only setting—is where most of our data gets corrupted or lost; MemMark allows us to extract all that information reliably.

Lu: The authors show this by using a distribution-preserving sampler that can handle the semantic realization carrier, which is one of the highest capacity options they found.

Meng: I’m looking at the performance metrics and it seems like we're talking about one point one six to one point two six bits of usable entropy per decision for different carriers, which is a lot of information packed into a single memory update.

Lalam: The ability to reliably recover data from the snapshot means that AI agents can build knowledge over long periods with confidence, and Lalam believes this is crucial for developing trust in autonomous systems.

The Conclusion: Tom: As we wrap up our discussion on MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems, the authors have demonstrated that durable attribution is possible even when logs are lost or corrupted.

Jane: It’s a robust way to prove who wrote a memory entry by tying the signature to the backend's own choices, not to a separate piece of metadata that could be edited.

Lu: The fact that it works across different backends like A-MEM and G-RAPHITI shows incredible generality for state-evolutionary systems.

Meng: From an engineering view, this means we can deploy these watermarking strategies in heterogeneous environments without needing a specific backend to support the attribution natively.

Lalam: Lalam feels that MemMark's ability to achieve trustworthy, verifiable memory is exactly what the future of AI needs to build reliable and consistent agents.

Tom: And with those insights, it's time for us to wrap up our discussion on this fascinating paper and look forward to the next one.

Conclusion: Tom: So, wrapping up our deep dive into "MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems," it really feels like we’ve seen a significant step forward in making AI memory traceable and accountable.

Jane: Exactly, Tom. What's so powerful about this paper is that it addresses the 'black box' problem of agent memory—it gives us tools to know *when* and *how* an AI decided something was important enough to remember.

Lu: I mean, thinking about the sheer complexity of long-term memory in a sophisticated agent, knowing that attribution watermark exists fundamentally changes how we view computational history.

Meng: And from an implementation standpoint, having that attribution signal means we can build systems that are auditable, which is huge if you think about enterprise integration or critical infrastructure.

Lalam: It’s more than just a technical fix; it’s an architectural shift toward transparency, making advanced AI more trustworthy and culturally integrated into human decision-making processes.

Tom: Trustworthiness—that's the keyword here, Jane. We went from discussing just memory capacity to discussing memory *integrity* and *provenance*.

Jane: It really shifts the focus from 'how much can it remember?' to 'can we trust what it remembers, and can we prove where that information came from?'

Meng: If we could guarantee that the memories an agent relies on actually came from the right source at the right time, that opens up possibilities for fields like medical diagnostics or complex regulatory compliance.

Lu: I just love thinking about applying this concept to historical simulations—being able to watermark which specific 'memory state' influenced a simulated outcome years later.

Tom: It’s wild how much this one paper touches on everything from engineering reliability to deep philosophical questions about agency itself.

Jane: Absolutely, it makes us think that the future of AI isn't just bigger models, but smarter, more accountable ones.

Lalam: And by establishing clear attribution standards like these, we help build a foundation where advanced AI can augment human culture without eroding trust or accountability.

Meng: We're definitely going to keep our eyes peeled for how companies actually try to scale this out of the lab and into real-world products.

Tom: Alright, team, we’ve got some incredible insights to take away from "MemMark." Thanks so much for joining us today!

cs.CR

Submitted: 2026-05-24

Updated: 2026-08-25

Importance score: 79/100

The gist: The paper focuses on "MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems," detailing a comprehensive evaluation of attribution watermarking techniques against

Key concepts

Attribution Watermarking
MemMark embeds a secret signal within the LLM's decision-making process for updating memory items. This ensures verifiable provenance by tying the signature to the backend's own choices rather than relying on external metadata that could be edited.
State-Evolutionary Systems
These are complex AI agents that build knowledge over long periods, often requiring long-term memory storage. MemMark addresses the issue of data corruption or loss in these systems by allowing reliable recovery from the final memory snapshot.
Full Payload Recovery
This is the technical ability to extract all stored information from a single, final memory snapshot. MemMark enables this reliable extraction even when external logs are missing, which is crucial for building trust in autonomous systems.

Terminology

Summary

The paper focuses on MemMark: State-Evolution Attribution Watermarking for Agent Long-Term Memory Systems, detailing a comprehensive evaluation of attribution watermarking techniques against various forms of memory tampering and data corruption. The methodology involves defining specific attack vectors, measuring recovery rates, and comparing performance across multiple large language models (LLMs).

The evaluation utilizes nine distinct memory-lifecycle attacks, grouped by the type of operation performed on the audit record. These attacks simulate malicious attempts to tamper with committed records:

1. Content-tamper attacks — modify committed record bytes in place:

These attacks involve modifying existing records while leaving other fields and the leaf set intact. Specific examples include:

  • para. rewrite: Silently mutating the probabilities field of a commitment fail record.

  • supersession: Replacing a selected candidate id with a sibling commitment fail candidate, simulating an older fact being overwritten by newer information.

  • edge relabel: Appending [RELABEL] to the selected candidate’s commitment fail (KG) payload.text, simulating relabeling of an entity–relation edge in the knowledge graph.

  • subgraph reanch: Appending [REANCHOR] to ctx t and rotating the commitment fail candidate list, simulating a change of root anchor for a subgraph.

  • Manual edits also cover generic record tampering, RAG-WM paraphrase attacks, Graphiti native factinvalidation chains, and KGMark edge perturbation or anchor swap.

2. Leaf-removal attacks — remove authenticated leaves from the trace:

These attacks simulate the deletion or loss of authenticated information:

  • pruning: Deleting a fraction of audit-record leaves uniformly at random, controlled by the attack strength.

  • dedup: Finding records duplicated by selected payload missing leaves text and removing secondary copies, retaining only the canonical leaf.

3. Synthesis/restructuring attacks — rewrite or add records without valid openings:

These attacks involve injecting or collapsing records:

  • compaction: Collapsing the candidate set by removing one commitment fail candidate, simulating multiple memories being merged into a single summary.

  • poisoning: Injecting fabricated audit records into the leaf set (additive, not deletive).

The paper reports performance using several metrics to quantify attribution signal retention:

1. Watermark Performance (Table 8):

Performance is measured via F1 scores for both watermarked (wm) and non-watermarked (nwm) scenarios, specifically for Qwen3.6-flash across ten LoCoMo samples. The watermark metrics use the A-M EM backend.

2. Attack Recovery Metrics (Table 10):

For the specialized RQ4 evaluation, two key metrics are used:

  • Rec: Post-attack bit recovery.

  • WK: Defined as Rec - WrongKey. The significance of this metric is that Positive WK means the attacked snapshot retains more attribution signal than an incorrect key.

The models evaluated include Deepseek-V4-pro, Qwen3.6-flash, and GLM-5. These models are tested across various attack families (Content 1–5, Removal 1–2, and Synthesis 1–2) at specified attack strengths s in 0.1, 0.3, 0.5.

The comparative results demonstrate the performance of the models across the defined attacks:

  • Content-tamper Attacks (Content-1 through Content-5): The models show varying levels of recovery and WK across these five content tampering scenarios. For example, Deepseek-V4-pro achieves a WK of +0.15 in Content-1, +0.20 in Content-2, +0.15 in Content-3, +0.10 in Content-4, and +0.12 in Content-5 (at the s=0.1 strength).

  • Leaf Removal Attacks (Removal-1 and Removal-2): The models maintain high recovery rates for these attacks; for instance, Deepseek-V4-pro achieves a WK of +0.70 in both Removal-1 and Removal-2 at s=0.1.

  • Synthesis Attacks (Synth-1 and Synth-2): Performance remains strong across synthesis attacks, with Deepseek-V4-pro achieving WK values such as +0.42 (Content s=0.1, Synth-1) and +0.62 (Content s=0.1, Synth-2).

In summary, the paper provides a rigorous framework for testing attribution watermarking by simulating nine specific memory lifecycle attacks, quantifying model

Improvements for AI systems

The research presented details critical vulnerabilities in the attribution and verifiability of LLM outputs when subjected to targeted data manipulation. To mitigate these risks and achieve enterprise-grade trustworthiness, I recommend implementing a multi-layered, cryptographically secured audit architecture.

Improvement: The core output generation process must be fundamentally separated from the final natural language rendering layer. All factual claims and derived relationships must first be constructed within a formalized, immutable Knowledge Graph (KG) structure—the Audit Record. This KG should utilize Merkle tree structures to commit to the state of facts at every step of the reasoning chain.

What the Improved System Can Do:

  • Guaranteed Traceability: Every claim (e.g., The CEO is John Doe) must link directly back to a specific, verifiable edge or node within the committed KG structure, eliminating reliance on merely citing surrounding text.

  • Differential Fact-Checking: The system can automatically detect subtle shifts in factual relationships (e.g., identifying if an entity's role has changed from founder to advisor) by comparing the current commit state against historical commits, mimicking the supersession and edge relabel low-level signals.

  • Low-Level Signature Detection: Implement specific parsers to detect markers associated with structural tampering, such as synthetic markers like [PARAPHRASE], [REANCHOR], or [RELABEL] inserted into the audit record.

  • Structural Integrity Checks (Pruning/Deduplication): The system must maintain a canonical set of facts. If an input trace is missing leaves (pruned) or contains redundant entries, the verifier must flag this loss of information and quantify the potential loss of attribution signal (WK), forcing human review before output generation.

  • Compaction/Poisoning Detection: The system must enforce a single source of truth rule. If multiple distinct facts are merged into a single summary statement (compaction), the verifier must break the statement apart and require explicit, verifiable commitments for each individual component, rather than allowing an aggregated claim.


The resulting AI system moves beyond being a mere text generator and becomes an Auditable Reasoning Engine (ARE).

Capability Functionality Provided Business Impact

:---:---:---

Verifiable Output Every generated sentence is linked to a committed, immutable set of facts (KG). The system can prove how it arrived at the conclusion. Eliminates hallucination risk in high-stakes environments (medical, legal, financial).

Tamper-Proof Audit Trail The commitment layer records the entire state transition. Attempts to modify facts or remove supporting evidence are mathematically detectable and flagged instantly. Meets stringent regulatory compliance requirements (e.g., HIPAA, GDPR) by providing undeniable proof of data handling integrity.

Resilience Scoring (WK) Provides a quantitative score indicating the degree of information loss or structural assumption made during generation, forcing user caution when the confidence is low. Manages risk proactively by defining clear operational boundaries for AI use cases, saving millions in potential liability.

Sources

Related papers