MaSRead: Content-Addressed Reading of Replicated Latent Stores
Carlos Baquero, Luís Brito, João Resende
University of Porto · Polytechnic Institute of Viana do Castelo · University of Porto
cs.AI, cs.LG, cs.MA
Submitted: 2026-07-21
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
The gist: The paper addresses the problem of reading replicated latent stores in multi-agent systems where "independent agents that reason in latent space can share computed state as key–value cache
Terminology
Summary
The paper addresses the problem of reading replicated latent stores in multi-agent systems where independent agents that reason in latent space can share computed state as key–value cache fragments rather than text.
These fragments are merged by a conflict-free replicated data type
to form a store that converges under any delivery order or duplication.
The central challenge is that a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability.
The setting involves "language-model agents reason and communicate in latent space rather than in text: instead of exchanging natural-language messages, each agent encodes its input into the transformer's key–value (KV) cache and passes that latent state on. When
many agents contribute, their latent states accumulate into a shared store, a replicated collection of KV fragments, each the distilled product of one agent's reasoning over its own input. The motivating scenario is that
an agent monitoring a long stream of routine events need not forward the stream; it could instead reason over it and contribute only the latent of the one anomaly it found, so that the store holds distilled reasoning rather than raw text."
The paper demonstrates that the obvious approach is to colocate the fragments, laying their caches side by side and reading the concatenation
fails. The authors show this fails. Colocated fragments interfere: a read aimed at one fragment is corrupted by the others, and the corruption worsens as the store grows.
The experimental evidence shows that at k = 2 the colocated read is already unreliable, at 0.43; by k = 8 it has collapsed to 0.00.
However, the fragments are not lost: the same store, read one fragment at a time under a mask to an oracle-designated target block, answers every query, at 0.99 or above for every k.
The paper concludes: The information is present in the store, but reading it whole does not give reliable selective access to it, and the loss grows with the store.
The failure is characterized as interference, not loss
: Reading over the colocated cache, a read aimed at one fragment is not confined to it: the answer it commits to is drawn from the wrong fragment or fused across several.
The paper notes this is sharpest when the fragments are lexically similar
and when the store is large.
The paper's central contribution is the masked signature read (MaSRead)
which routes through opaque keyed tag sets derived from fragment words and decodes each selected fragment under a hard attention mask that hides the rest.
The mechanism works in three operations:
-
Content-derived signatures: "A writer normalizes its fragment text locally: lowercase alphabetic content words, simple stemming, function-word removal, and no digits. Under a store-scoped secret key K, each normalized word w becomes the first 128 bits of a domain-separated HMAC-SHA256(K, d ∥ w).
The writer
sends the resulting enumerable tag set with the cache and may then discard the source text." -
Routing via graph walk:
A query is tagged into the same domain, and the read routes to the fragments whose signatures its tags select, preferring the largest overlap.
The algorithm "seeds a frontier with the query's tags and matches them, by set membership, against the fragments' signatures; each matched fragment is visited, and its own signature tags are unioned into the frontier, so the search expands without ever enumerating a Bloom filter." -
Masked decoding: "Having located a fragment, the read decodes over the render under a hard attention mask that admits only that fragment's block and hides every other, so the interference of Section 3 cannot arise: the model attends to one fragment and reads it as if it stood alone."
The mechanism is validated by a control: "a read under the wrong mask, admitting the partner's block instead of the target's, scores 0.000, and its answers are the partner's value. This is what makes the repair addressing rather than denoising. The mask does not clean up a noisy read; it selects which fragment is read."
The paper defines a replicated latent-store read contract
with six separable obligations:
-
Storage and convergence:
preserve the same immutable fragments at every replica, but do not make any one fragment selectively readable
-
Routing:
must find the fragments required by a later query; our opaque signature walk does so only when a lexical path connects the query to them
-
Addressing:
must then expose the selected fragment without exposing its neighbors
-
Recovery:
asks whether the selected cache can be decoded into the content it holds, a model-dependent step measured by fragment restatement
-
Isolation:
asks whether unrelated stored fragments can alter that recovery
-
Composition:
asks whether the reader can turn the recovered fragments into the final answer
The paper states: "This decomposition makes an end-to-end score interpretable and is the paper's main general lesson. Failure can mean that evidence was absent, missed by routing, misaddressed, decoded incorrectly, contaminated by other fragments, or recovered but not composed; the remedies differ."
The store is formally defined: "Each agent, having encoded its input query-blind under one shared, frozen model, contributes a fragment: the key–value cache of that encode and an immutable lexical-addressing sidecar, named together by a content identifier. The state
is the set of fragments it has received and
the merge is set union."
The paper proves convergence: Set union is commutative, associative, and idempotent, so every delivery history with the same delivered elements yields the same S.
The rendering is a deterministic function of the set, the render is byte-identical across delivery orders.
The paper evaluates on synthetic stores of four families
: a chain of fictional-unit conversions,
an affine pipeline of machines,
a symmetric
constraint system, and a hub
with a single dense fragment indexed against two small tables.
Results show: The read reaches the full-text ceiling on three of the four families: the chain, the pipeline, and the symmetric store all land at 0.97 or above.
The hub is the exception at 0.44, where coverage and the hub's multi-value readout are both 1.00, so the store returns every fact the query needs. What the frozen reader then fails is composing those facts into the answer.
The paper tests whether the read survives this: we take a store that reads well, merge in unrelated fragments drawn from a disjoint family that shares no words with the query, and read for the original target as the store fills.
Results show: "The masked read holds near 0.90 across the whole sweep, undiminished as the store fills with unrelated fragments to twenty times what a query needs. The colocated read over the same store collapses, from 0.94 to near zero by D = 16."
On two standard multi-hop question-answering sets: MuSiQue 2-hop and HotpotQA bridge,
the paper finds: With eight topical distractors the masked read holds, 0.44 on MuSiQue and 0.57 on HotpotQA, while both unaddressed reads collapse to near 0.03.
The paper notes: On a clean store the masked read does not win: with no distractors to interfere, both unaddressed reads are as good or better.
The paper cuts the confound by varying the agent's latent budget l, the number of latent reasoning steps it takes while encoding a fragment, down to l = 0 where it does none.
Results show: "Three of the four structures are flat: the chain holds 0.93 at every l, the pipeline and the symmetric store stay at or above 0.99. Removing latent computation entirely costs nothing. The read recovers the encoded fact; it does not depend on the agent having reasoned over it."
The hub decomposition shows: Coverage, restatement, and hub extraction are all 1.00, and the decoded facts are parseable at 0.99: the read returns every value the answer needs, in usable form.
However, "A symbolic composer, a fixed rule applied to those decoded facts instead of the model, answers at 0.99. Nothing the answer requires is missing from the read; the read's output already determines the answer. Yet the frozen reader, handed the same decoded facts, answers at 0.40."
The paper provides a cost equation: Ttotal (q, S) = Troute (q, S) + Σ Tread (Lf) + Tcompose (V (q, S)).
It states: Only the reading of a single located fragment is independent of the store. The full read is not: the walk scans the signatures and visits every fragment that shares a frontier tag, so Troute generally grows with S.
However, "What does not grow is the cost of reading one already located fragment... the target block is extracted and decoded at the native positions it was encoded at, so its cost is the block's own length O(Lf), independent of the store size S."
The paper is explicit about limitations:
-
Lexical routing: "A query reaches a fragment only through shared lexical terms (represented as equal tags), so a needed fragment that shares no routing token is never read: on the DISCOUNT family this happens on every item, and a colocated read that keeps everything in view does better there than the addressed read."
-
Composition bounded by reader:
The store delivers the facts a query needs when routing reaches them; composing them into an answer is bounded by the reader.
-
Not purely latent:
Writers and authorized query clients normalize text before mapping its words to opaque tags, and the reader recovers each fragment by decoding a short text restatement before composing.
-
Not constant-time search:
routing scans store metadata, V may grow with lexical connectivity, and every visited fragment is recovered and composed. Only the materialized read of one already located block is independent of S.
The paper states: "The read we study is not text retrieval, and we do not benchmark against it. A pipeline that stores and re-reads text is a different regime: it presupposes that the source text is retained, which the setting here does not provide. The paper claims
addressability, robustness to contamination, and read cost, not superiority over text on those axes."
The paper's central claim: "Once agents reason in latent space, the store they produce can be read for queries they never saw, when those queries are connected by content to the fragments they need, provided the store is addressed rather than merely colocated. Addressing, not merging, is the operation that makes a replicated latent store usable, and it is the operation this paper supplies."
Improvements for AI systems
Based on the paper, here are specific improvements I can implement in AI systems:
Implementation: Add a routing layer that tags KV-cache fragments with opaque HMAC signatures derived from content words, then reads fragments under hard attention masks instead of decoding concatenated caches.
What the improved system can do: When multiple agents contribute KV fragments to a shared store, a later query can selectively retrieve and decode the specific fragment it needs, even when the store grows to 64+ unrelated fragments. Currently, naive concatenation causes interference that collapses accuracy from 0.94 to near zero by 16 fragments; MaSRead maintains 0.90 accuracy at 64 fragments.
Abstract
Independent agents that reason in latent space can share computed state as key-value cache fragments rather than text. Merged by a conflict-free replicated data type, these fragments form a store that converges under any delivery order or duplication. Yet a later query, unknown at encode time, cannot reliably read the merged cache: colocated fragments interfere, so colocation is not addressability. MaSRead addresses the read to content. It routes through opaque keyed tag sets derived from fragment words and decodes each selected fragment under a hard attention mask that hides the rest. Under lexical connectivity, a graph walk reaches the fragments required by a multi-hop query. Across chain, pipeline, symmetric, hub, and natural-language stores, MaSRead recovers visited fragments in isolation, remains effective as unrelated fragments accumulate, and transfers to another model family. After routing, materialized decoding depends on fragment length rather than total store size; end-to-end work still includes store-dependent routing and one read per visited fragment. The limits are explicit: lexical routing can miss disconnected evidence, and answer composition remains bounded by the frozen reader. Thus a replicated latent store becomes selectively readable for later queries when the needed fragments connect to the query through content.
Sources
- Cache Merging as a Convergent Replicated State for Multi-Agent Latent Reasoning
- Enabling Agents to Communicate Entirely in Latent Space
- LatentMem: Customizing Latent Memory for Multi-Agent Systems
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Conflict-Free Replicated Data Types for Neural Network Model Merging: A Two-Layer Architecture Enabling CRDT-Compliant Model Merging Across 26 Strategies
- Prompt Cache: Modular Attention Reuse for Low-Latency Inference
- The Llama 3 Herd of Models
- CLaRa: Bridging Retrieval and Generation with Continuous Latent Reasoning
- FlashMem: Distilling Intrinsic Latent Memory via Computation Reuse
- Agent Primitives: Reusable Latent Building Blocks for Multi-Agent Systems
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- SnapKV: LLM Knows What You are Looking for Before Generation
- Deliberation in Latent Space via Differentiable Cache Augmentation
- Block-Attention for Efficient Prefilling
- Conflict-free Replicated Data Types (CRDTs)
- ProphetKV: User-Query-Driven Selective Recomputation for Efficient KV Cache Reuse in Retrieval-Augmented Generation
- Model Tells You Where to Merge: Adaptive KV Cache Merging for LLMs on Long-Context Tasks
- Efficient Streaming Language Models with Attention Sinks
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection