Frozen Memory Is Not Enough: Rethinking External Memory as Extraction

arXiv:2608.17050 · cs.CL, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Frozen Memory Is Not Enough".

Jane: The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re talking about this paper by Mingyuan Li, Guangsheng Yu, Xu Wang, and Shaoxiong Ji from the ELLIS Institute in Finland and University of Technology Sydney. The title itself really sets the stage for what they are investigating: whether frozen memory is truly sufficient on its own when you try to transfer it between different AI backbones.

Jane: Exactly, Tom. They are looking at this Engram-style hashed external memory, which stores learned information in an external table that a small neural interface consumes rather than raw documents being inserted into the context window.

Lu: The authors pose the central question of whether you can move that frozen memory across different backbones and still get a useful signal out, or if it’s just model-specific parameterization. They suggest the portability requires an operational test: removing the source backbone and seeing if another model can still extract a useful signal from that table.

Meng: That sounds like a rigorous test, but I wonder how they handle the alignment between models that might have completely different tokenizers. If Model A and Model B use different ways of breaking down text, how do you ensure the memory address maps correctly?

Lalam: That tokenizer issue is something I’ve been thinking about because if the addressing isn't standardized, then transferring knowledge across models becomes impossible to do reliably.

The paper's summary: Tom: Moving on to what they actually found in this study, the authors present an evidence chain demonstrating that the design of the target-side reader is first-order—meaning it’s critical for performance. They show consistent gains across diverse model families and settings when they adjust how much memory is injected and how many branches the reader has.

Jane: It seems they pinpoint a few key levers: you can change where in the backbone you inject the memory, like single-layer versus dual-layer injection, or increase the number of reader branches from one up to four.

Lu: That’s really insightful because it shows that stronger target-side readers drive better performance and achieve a new state of the art. They found that increasing branches from R=one to R=four raises average QA accuracy from thirty-seven point five to thirty-eight point five, which suggests multi-branch gating provides complementary key-gating pathways over the same shared value representation.

Meng: So, it’s not just about having a memory table; it’s about designing an interface that lets the target model actually read and use that information in a way that maximizes its potential. That makes sense from an implementation perspective.

Lalam: It really shifts the focus from just storing data to designing the access mechanism, which is something we can definitely work on improving for better cultural integration across different platforms.

The paper's improvements: Tom: The authors are suggesting several ways to improve this process, focusing heavily on standardizing the address space through a "tokenizer-agnostic canonicalization pipeline." This involves using functions like NFKC normalization and deterministic N-gram hashing to map text sequences to fixed memory indices in a shared table.

Jane: That standardization is crucial for portability because it ensures that the same piece of text gets mapped to the exact same memory slot, no matter which tokenizer the target model uses. That’s a big step toward making external memory truly cross-model usable.

Lu: They are proposing this protocol as a way to test if frozen Engram-style hashed external memory is reusable or merely co-adapted, and they show how this standardized addressing enables that operational test.

Meng: From an implementation viewpoint, I see the canonicalization pipeline as a necessary preprocessing step to make the memory table address space consistent across different model families. It’s like creating a universal dictionary for the knowledge base.

Lalam: If we can implement that kind of robust addressing system, it means our systems won't be locked into proprietary text formats or specific model vocabularies, which opens up so much flexibility in how we share and reuse knowledge.

Conclusion: Tom: So, to wrap up this discussion on "Frozen Memory Is Not Enough: Rethinking External Memory as Extraction," the main conclusion is that cross-model frozen-memory extraction turns external memory portability into an evaluation framework. The key finding is that downstream gains cannot be explained by parameter count alone; you need meaningful addressing and a sufficiently expressive reader interface to get successful transfer.

Jane: I agree, Tom. The paper shows that the success of transferring learned information depends on whether a new backbone can address it, align to it, and extract useful signals without needing to retrain the stored table.

Lu: The most significant part is that reader design is shown to be first-order: consistent gains across diverse model families and settings are observed when you optimize the reader’s placement and capacity. That tells us exactly where to focus our architectural efforts for future memory systems.

Meng: From my side, it confirms that target-data efficiency is high; you can reach strong intrinsic performance with fewer target-side updates than learning a new memory from scratch, which is great for deployment costs.

Lalam: I just think this entire study reinforces the idea that knowledge isn't just about having data stored somewhere; it's about the intelligent interface we build to read and utilize that data effectively across different systems.

ELLIS Institute Finland · University of Turku · University of Technology Sydney

cs.CL, cs.AI

Submitted: 2026-08-17

Updated: 2026-09-30

Code: https://github.com/OLAResearch/XMemTransfer

Project page: https://olaresearch.org/XMemTransfer

Importance score: 85/100

The gist: The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models, specifically focusing on whether frozen memory alone is sufficient

Key concepts

Portability vs. Co-adaptation
This concept distinguishes whether frozen memory contains genuinely reusable knowledge or just information tailored specifically to the original model's structure. The paper argues that for memory to be portable, it must contain information beyond what is unique to the source model's specific architecture.
Tokenizer-Agnostic Canonicalization
This is a standardized process used to map different text inputs from various models into a single, consistent format before hashing. It involves cleaning and normalizing raw tokens (like lowercasing) and then using deterministic functions to create fixed memory indices, allowing the same piece of content to be found regardless of the tokenizer.
Target-Side Reader Design
This refers to the architectural choices made on the receiving model that determines how it accesses and extracts meaning from the frozen memory. Choices like where memory is injected or how many retrieval branches are used directly influence whether useful knowledge can be successfully pulled out.

Terminology

Summary

The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models, specifically focusing on whether frozen memory alone is sufficient or if it requires a compatible target-side reader for successful knowledge extraction. This research addresses a fundamental question in model architecture design: when moving learned information between backbones, does the frozen artifact retain its utility, or is its value contingent upon the target model's ability to read and integrate it? The findings suggest that while frozen memory is portable across various model families and scales, successful reuse critically depends on the interface through which a target backbone addresses and consumes the stored representations.

The Core Hypothesis: Portability vs. Co-adaptation

The study frames cross-model frozen-memory transfer as an operational test to determine if a memory table contains usable information beyond source-specific co-adaptation. The central claim is that for portable hashed external memory, memory presence alone is not enough; successful reuse depends on the interface through which a target backbone addresses and consumes the stored representations. This distinction separates a reusable knowledge substrate from mere model-specific parameterization.

The Transfer Protocol: Standardizing Address Space

To enable transfer across different tokenizers, the method standardizes addressing through a tokenizer-agnostic canonicalization pipeline. The process involves:

  1. Canonicalizing raw tokens using a function P (NFKC normalization, lowercasing, accent stripping).

  2. Performing N-gram Hashing using deterministic functions ϕn,k to map canonicalized sequences to fixed memory indices in the shared memory table E.

  3. Freezing the source Engram memory table EA and training only a lightweight target-side reader W(B) for extraction.

The Target-Side Reader Design: The Determinant of Performance

The performance ceiling is determined by the reader's design, which is treated as a first-order factor. Key architectural choices include:

  1. Memory Placement: Varying the layers (e.g., single-layer vs. dual-layer injection) at which memory is injected into the target backbone.

  2. Retrieval Capacity: Increasing the number of reader branches (R), such as moving from R=1 to R=4, which increases target-side alignment capacity without increasing memory size.

  3. Reader Structure: Utilizing a shared-value reader family where different branches operate on the same retrieved content but differ in their key-gating pathways.

Evidence Chain: Reader Design Drives Gains

The evidence chain demonstrates that reader design is critical, showing that consistent gains across diverse model families and settings are observed. Specifically:

)&Reader Placement:

**) Dual-layer injection (e.g., layers 2 and 10) increases average QA accuracy from 34.2 to 37.5 on LLaMA models, suggesting multiple integration opportunities are more beneficial than a single late-layer interface. **

) Reader Capacity:

**) Increasing branches from R=1 to R=4 raises average QA accuracy from 37.5 to 38.5, demonstrating that multi-branch gating provides complementary key-gating pathways over the same shared value representation. **

Downstream Utility and Limitations

The utility of transferred memory is task-dependent:

  1. Downstream Gains: Transferred memory shows selective rather than universal benefits; it improves factual and evidence-related tasks like RTE, SciQ, and BoolQ, but has little effect on broader reading comprehension.

  2. TruthfulQA Boundary: Performance decreases on TruthfulQA across target scales (ranging from −0.2 to −0.8 accuracy points), suggesting that transferred memory is most beneficial when the task rewards access to factual or evidence-related information, but it may be neutral or mildly harmful when success depends on calibration.

  3. Efficiency: Frozen-memory transfer is target-data efficient, reaching strong intrinsic performance with substantially fewer target-side updates than learning a new memory from scratch, though asymptotic ceilings are not necessarily higher than scratch training.

Conclusion: The Evaluation Framework

The paper concludes that cross-model frozen-memory extraction turns external memory portability into an evaluation framework. A memory artifact should be judged by whether a new backbone can address it, align to it, and extract useful signal without retraining the stored table. The main takeaway is that the downstream gains cannot be explained by parameter count alone: successful transfer depends on meaningful memory addressing and a sufficiently expressive reader interface. Furthermore, the cost analysis shows that the one-time provider cost is amortized after approximately nine adapted consumers.

Key Takeaways:

) Portability depends not only on the frozen table but also on the target-side reader.

**) Reader design is a "first-order determinant of transfer quality.

Improvements for AI systems

Based on the provided scientific paper, here are specific, actionable improvements to AI systems derived from the cross-model frozen-memory transfer (Engram) framework:


The core improvement is shifting knowledge utilization from brittle, model-specific weights to a reusable, addressable external knowledge artifact. This enables knowledge portability across different model backbones and architectures.

Here are the specific improvements and capabilities:

  1. Cross-Model Knowledge Transfer without Fine-Tuning (Zero Target Training):

  2. Mechanism: Instead of retraining a target model on a new corpus to learn facts, you can transfer frozen memory from a source model (e.g., LLaMA) to a target model (e.g., Mistral). The only training required for the target is fitting a small, lightweight reader.

  3. Capability: This allows deploying knowledge learned by Model A into Model B instantly, bypassing the massive computational cost and data requirements of full fine-tuning for every new task or model variant.

  4. Robustness to Tokenizer Mismatches:

  5. Mechanism: The system uses a tokenizer-agnostic canonicalization pipeline (NFKC normalization, lowercasing, accent stripping) before hashing memory addresses. This ensures that the same piece of text maps to the same memory slot regardless of whether Model A and Model B use different tokenizers (e.g., GPT vs. LLaMA).

  6. Capability: The transferred knowledge remains accessible across diverse model families and tokenizer boundaries, solving a major bottleneck in cross-model reuse that current RAG or adapter methods struggle with when architectures diverge significantly.

  7. Optimized Knowledge Extraction via Reader Architecture Design:

  8. Mechanism: The target system can be designed with a multi-layer, multi-branch reader (e.g., dual-layer injection at layers 2 and 10 with R=4 branches). This reader is optimized to intelligently gate and integrate the retrieved memory vector into the target model's hidden state.

  9. Capability: The AI system gains significantly higher accuracy (up to a 19.9% improvement on QA tasks) by learning the optimal way to interpret and route external knowledge, effectively creating a highly specialized knowledge extraction module tailored for the specific target backbone's residual stream.

  10. Task-Specific Knowledge Specialization:

  11. Mechanism: The system can be tuned for specific downstream goals (e.g., high factual retrieval vs. resistance to misinformation) by controlling the training corpus used during Phase 2 reader adaptation (e.g., training on a QA-heavy corpus like Nemo HQ-DQA).

  12. Capability: This allows the transferred memory to exhibit superior performance on tasks where factual evidence is key (like NQ, WebQA, and SciQ) while maintaining neutral or even slightly negative performance on tasks requiring calibrated truthfulness (like TruthfulQA), providing a fine-grained control over knowledge utility.

  13. Data-Efficient Initialization:

  14. Mechanism: The transferred memory provides a superior structured initialization compared to training a new memory from scratch, reaching competitive performance with significantly fewer target-side adaptation tokens (e.g., 5M tokens vs. 34M parameters).

  15. Capability: This drastically reduces the required data budget for adapting a new model to utilize existing knowledge, making knowledge reuse highly data-efficient and accelerating model deployment cycles.

  16. Dynamic Knowledge Management:

  17. Mechanism: The external memory is addressable and explicit; entries can be inspected, removed, or replaced without ever touching the backbone weights.

  18. Capability: This enables modular knowledge updates—you can swap out specific facts or knowledge modules to improve performance without requiring a full model retraining cycle, offering unprecedented auditability and flexibility in managing learned information.

Sources

Related papers