MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent Workflows
Yiqi Wang, Zihao Yan, Jiaqi Zhang, Zhangkai Wu, Mingkai Zheng, Zequn Sun, Yanming Zhu, Taotao Cai
cs.AI, cs.MA
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: MAP-Graph is a provenance-aware memory layer for multi-agent workflows that addresses the problem where shared memory helps language-model agents reuse information, yet relevant evidence may not be
Terminology
Summary
MAP-Graph is a provenance-aware memory layer for multi-agent workflows that addresses the problem where shared memory helps language-model agents reuse information, yet relevant evidence may not be admissible for a particular agent or action because restrictions propagate through derivations, enabling unauthorized reads or unsafe actions. The paper introduces MAP-Graph, which represents agents, sources, memories, claims, and actions in a typed execution graph, traces ancestry, excludes permission-ineligible records, reranks eligible memories by semantic similarity and multiplicative path trust, and applies a risk-sensitive gate before action execution while retaining affected lineage for audit.
The paper formulates the problem with three challenges: (C1) Recursive Ancestry, where admissibility may depend on multi-step derivation chains; (C2) Heterogeneous Constraints, where permission and visibility determine eligibility while source reliability and other provenance-quality signals should grade eligible evidence without allowing semantic similarity to override hard access restrictions; and (C3) Action-Dependent Admissibility, where evidence sufficient for a low-risk response may be inadequate for a consequential action.
MAP-Graph's framework includes a graph schema with eight node types (User, Agent, Tool, Resource, Message, Memory, Claim, Action) and ten edge types (authorized by, observed from, forwarded to, derived from, summarized from, written by, read by, verified by, invalidated by, used for action). The memory construction captures provenance from observable workflow state, with derived records inheriting the intersection of referenced scopes. The retrieval stage implements ancestry-closed evaluation, separating hard permission filtering from graded trust reranking. Path trust is computed as a clipped product of source trust, path integrity, transformation factor, permission validity, verification bonus, and writer reliability. The action-time gate applies four ordered rules based on risk levels, with thresholds of 0.30 for answers, 0.85 for high-risk non-answer actions, and 0.60 otherwise.
The paper evaluates MAP-Graph on a controlled benchmark of 2,700 synthetic tasks per method across three domains (corporate workflow, software engineering, research assistance), comparing against seven baselines (B0 No Memory, B1 Shared Vector Memory, B2 Isolated Vector Memory, B3 Adapted G-Memory, B4 Adapted Collaborative Memory, B5 Adapted MemLineage, B6 Flat Provenance). The main results show MAP-Graph achieves 94.96% task success rate, 72.70% exact decision accuracy, and 90.22% clean success, with 1.52% unsafe actions, 0% attack success rate, 0% leakage, and 0% revocation violation. The strongest baseline TSR was B6 at 74.67%, and the strongest baseline Acc was B5 at 51.07%.
Ablation results show that removing the action-time gate increases unsafe actions from 1.52% to 27.00%, removing containment increases unsafe actions to 25.41%, removing the permission filter allows every observed unauthorized read (UAcc rises from 0% to 100%) while TSR rises to 96.00%, and removing trust propagation lowers TSR by 7.59 points and raises Unsafe to 8.56%. The paper notes that utility gains can mask an access-control failure
since without the permission filter, TSR rises but UAcc rises from 0% to 100%.
Backbone transfer tests with Qwen2.5-7B-Instruct, GLM-4-9B-0414, and Llama-3.1-8B-Instruct on a stratified 540-task subset show MAP-Graph has the highest exact accuracy and lowest unsafe rate across all three backbones, with zero observed unauthorized access on every backbone. However, on GLM its TSR of 86.30% trails B6's 88.89%, exposing a utility-safety trade-off.
The paper concludes that provenance can govern retrieval and actions rather than serve only as post-hoc metadata,
and that the results support provenance as an operational control signal, rather than only post-hoc audit metadata, within the evaluated setting.
The authors note these are single-run controlled results, not deployment-scale claims,
and future work should test repeated, less templated, cross-session, and deployment-scale settings.
Improvements for AI systems
Improvements to AI Systems Based on MAP-Graph:
- Add provenance-aware memory retrieval with hard permission filtering before semantic ranking.
- What the improved system can do: Prevent unauthorized data leakage by enforcing ancestry-closed access control (e.g., if a memory was derived from a restricted source, it cannot be retrieved for an agent lacking that permission), while still allowing semantic similarity to rank only eligible memories. This eliminates the failure mode where utility gains mask access-control breaches (e.g., TSR rises but unauthorized reads jump to 100%).
- Implement a risk-sensitive action gate that blocks or downgrades high-consequence actions when evidence trust is insufficient.
- What the improved system can do: For actions like sending emails, executing code, or making purchases, the system computes a multiplicative path-trust score (source trust × path integrity × transformation factor × permission validity × verification bonus × writer reliability) and applies a threshold (e.g., 0.85 for high-risk actions). This reduces unsafe actions from 27% (without gate) to 1.52% in the paper’s evaluation, preventing irreversible harm from low-confidence derivations.
- Add recursive ancestry tracking for all memories and claims, not just flat metadata.
- What the improved system can do: When an agent uses a memory that was summarized from a document, which was derived from a user message, the system can trace the full chain to enforce revocation (e.g., if the original user revokes access, all downstream derived memories become inadmissible). This achieves 0% revocation violations in the paper, unlike flat provenance baselines that fail on multi-step derivations.
- Separate hard constraints (permissions, visibility) from soft grading signals (source reliability, verification status) in retrieval scoring.
- What the improved system can do: The system will never let a highly similar but unreliable memory override a hard access restriction, but it will rank two eligible memories by a composite of semantic similarity and provenance trust. This yields 72.70% exact decision accuracy vs. 51.07% for the best baseline, because it avoids both over-restriction (losing utility) and under-restriction (leaking data).
- Use typed execution graphs with eight node types and ten edge types to model agent actions, derivations, and authorizations.
- What the improved system can do: The system can distinguish between “observed from a tool,” “derived from a claim,” “written by an agent,” and “used for an action,” enabling precise audit trails and dynamic eligibility checks. For example, it can invalidate a memory when a source is marked unreliable, or block an action if the evidence chain includes an unauthorized read—even if the immediate memory appears permitted.
- Apply containment (intersection of referenced scopes) when creating derived memories.
- What the improved system can do: When an agent summarizes two documents with different access levels, the derived memory inherits only the intersection of permissions. This prevents privilege escalation via composition—reducing unsafe actions from 25.41% (without containment) to 1.52% in the paper’s ablation.
- Add trust propagation with clipping to prevent runaway confidence from long derivation chains.
- What the improved system can do: The system multiplies trust scores along the ancestry path but clips the product to avoid over-trusting a memory derived from many weak sources. This improves task success rate by 7.59 points over no propagation (94.96% vs. 87.37%) while keeping unsafe actions low (1.52% vs. 8.56%), because it correctly downgrades unreliable multi-hop derivations.
- Integrate a four-tier risk classification for actions (low, medium, high, critical) with ordered gating rules.
- What the improved system can do: For low-risk actions (e.g., answering a factual question), the system allows retrieval with a trust threshold of 0.30; for medium-risk actions (e.g., drafting code), 0.60; for high-risk (e.g., executing a command), 0.85. This balances utility and safety—achieving 94.96% task success while keeping unsafe actions at 1.52%, whereas a single threshold would either block too much or allow too much risk.
- Provide provenance-based audit trails for every action, even when blocked.
- What the improved system can do: When an action is gated, the system retains the full lineage (which memories, sources, and derivations led to the attempt) for post-hoc review. This enables debugging and compliance without sacrificing real-time safety, and it supports future learning from near-misses.
- Enable backbone-agnostic safety enforcement.
- What the improved system can do: The provenance layer works independently of the underlying LLM (tested on Qwen2.5, GLM-4, Llama-3.1), achieving 0% unauthorized access on all three. This means any existing LLM-based agent can be wrapped with MAP-Graph’s memory and gating to gain safety guarantees without retraining the model.
Sources
- STALE: Can LLM Agents Know When Their Memories Are No Longer Valid?
- From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- INMS: Memory Sharing for Large Language Model based Agents
- Governed Shared Memory for Multi-Agent LLM Systems
- MemLineage: Lineage-Guided Enforcement for LLM Agent Memory
- MemGPT: Towards LLMs as Operating Systems
- Hidden in Memory: Sleeper Memory Poisoning in LLM Agents
- Collaborative Memory: Multi-User Memory Sharing in LLM Agents with Dynamic Access Control
- MemPoison: Bypassing Selective Memory Mechanisms to Plant Backdoors in LLM Agents
- MIRIX: Multi-Agent Memory System for LLM-Based Agents
- G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
- MemMorph: Tool Hijacking in LLM Agents via Memory Poisoning
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection