Frozen Memory Is Not Enough: Rethinking External Memory as Extraction
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Frozen Memory Is Not Enough".
Jane: The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we’re talking about this paper by Mingyuan Li, Guangsheng Yu, Xu Wang, and Shaoxiong Ji from the ELLIS Institute in Finland and University of Technology Sydney. The title itself really sets the stage for what they are investigating: whether frozen memory is truly sufficient on its own when you try to transfer it between different AI backbones.
Jane: Exactly, Tom. They are looking at this Engram-style hashed external memory, which stores learned information in an external table that a small neural interface consumes rather than raw documents being inserted into the context window.
Lu: The authors pose the central question of whether you can move that frozen memory across different backbones and still get a useful signal out, or if it’s just model-specific parameterization. They suggest the portability requires an operational test: removing the source backbone and seeing if another model can still extract a useful signal from that table.
Meng: That sounds like a rigorous test, but I wonder how they handle the alignment between models that might have completely different tokenizers. If Model A and Model B use different ways of breaking down text, how do you ensure the memory address maps correctly?
Lalam: That tokenizer issue is something I’ve been thinking about because if the addressing isn't standardized, then transferring knowledge across models becomes impossible to do reliably.
The paper's summary: Tom: Moving on to what they actually found in this study, the authors present an evidence chain demonstrating that the design of the target-side reader is first-order—meaning it’s critical for performance. They show consistent gains across diverse model families and settings when they adjust how much memory is injected and how many branches the reader has.
Jane: It seems they pinpoint a few key levers: you can change where in the backbone you inject the memory, like single-layer versus dual-layer injection, or increase the number of reader branches from one up to four.
Lu: That’s really insightful because it shows that stronger target-side readers drive better performance and achieve a new state of the art. They found that increasing branches from R=one to R=four raises average QA accuracy from thirty-seven point five to thirty-eight point five, which suggests multi-branch gating provides complementary key-gating pathways over the same shared value representation.
Meng: So, it’s not just about having a memory table; it’s about designing an interface that lets the target model actually read and use that information in a way that maximizes its potential. That makes sense from an implementation perspective.
Lalam: It really shifts the focus from just storing data to designing the access mechanism, which is something we can definitely work on improving for better cultural integration across different platforms.
The paper's improvements: Tom: The authors are suggesting several ways to improve this process, focusing heavily on standardizing the address space through a "tokenizer-agnostic canonicalization pipeline." This involves using functions like NFKC normalization and deterministic N-gram hashing to map text sequences to fixed memory indices in a shared table.
Jane: That standardization is crucial for portability because it ensures that the same piece of text gets mapped to the exact same memory slot, no matter which tokenizer the target model uses. That’s a big step toward making external memory truly cross-model usable.
Lu: They are proposing this protocol as a way to test if frozen Engram-style hashed external memory is reusable or merely co-adapted, and they show how this standardized addressing enables that operational test.
Meng: From an implementation viewpoint, I see the canonicalization pipeline as a necessary preprocessing step to make the memory table address space consistent across different model families. It’s like creating a universal dictionary for the knowledge base.
Lalam: If we can implement that kind of robust addressing system, it means our systems won't be locked into proprietary text formats or specific model vocabularies, which opens up so much flexibility in how we share and reuse knowledge.
Conclusion: Tom: So, to wrap up this discussion on "Frozen Memory Is Not Enough: Rethinking External Memory as Extraction," the main conclusion is that cross-model frozen-memory extraction turns external memory portability into an evaluation framework. The key finding is that downstream gains cannot be explained by parameter count alone; you need meaningful addressing and a sufficiently expressive reader interface to get successful transfer.
Jane: I agree, Tom. The paper shows that the success of transferring learned information depends on whether a new backbone can address it, align to it, and extract useful signals without needing to retrain the stored table.
Lu: The most significant part is that reader design is shown to be first-order: consistent gains across diverse model families and settings are observed when you optimize the reader’s placement and capacity. That tells us exactly where to focus our architectural efforts for future memory systems.
Meng: From my side, it confirms that target-data efficiency is high; you can reach strong intrinsic performance with fewer target-side updates than learning a new memory from scratch, which is great for deployment costs.
Lalam: I just think this entire study reinforces the idea that knowledge isn't just about having data stored somewhere; it's about the intelligent interface we build to read and utilize that data effectively across different systems.
ELLIS Institute Finland · University of Turku · University of Technology Sydney
cs.CL, cs.AI
Submitted: 2026-08-17
Updated: 2026-09-30
Code: https://github.com/OLAResearch/XMemTransfer
Project page: https://olaresearch.org/XMemTransfer
Importance score: 85/100
The gist: The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models, specifically focusing on whether frozen memory alone is sufficient
Key concepts
- Portability vs. Co-adaptation
- This concept distinguishes whether frozen memory contains genuinely reusable knowledge or just information tailored specifically to the original model's structure. The paper argues that for memory to be portable, it must contain information beyond what is unique to the source model's specific architecture.
- Tokenizer-Agnostic Canonicalization
- This is a standardized process used to map different text inputs from various models into a single, consistent format before hashing. It involves cleaning and normalizing raw tokens (like lowercasing) and then using deterministic functions to create fixed memory indices, allowing the same piece of content to be found regardless of the tokenizer.
- Target-Side Reader Design
- This refers to the architectural choices made on the receiving model that determines how it accesses and extracts meaning from the frozen memory. Choices like where memory is injected or how many retrieval branches are used directly influence whether useful knowledge can be successfully pulled out.
Terminology
Summary
The paper investigates whether Engram-style hashed external memory remains useful and transferable across different language models, specifically focusing on whether frozen memory alone is sufficient or if it requires a compatible target-side reader for successful knowledge extraction. This research addresses a fundamental question in model architecture design: when moving learned information between backbones, does the frozen artifact retain its utility, or is its value contingent upon the target model's ability to read and integrate it? The findings suggest that while frozen memory is portable across various model families and scales, successful reuse critically depends on the interface through which a target backbone addresses and consumes the stored representations.
The Core Hypothesis: Portability vs. Co-adaptation
The study frames cross-model frozen-memory transfer as an operational test to determine if a memory table contains usable information beyond source-specific co-adaptation.
The central claim is that for portable hashed external memory, memory presence alone is not enough; successful reuse depends on the interface through which a target backbone addresses and consumes the stored representations.
This distinction separates a reusable knowledge substrate from mere model-specific parameterization.
The Transfer Protocol: Standardizing Address Space
To enable transfer across different tokenizers, the method standardizes addressing through a tokenizer-agnostic canonicalization pipeline.
The process involves:
-
Canonicalizing raw tokens using a function P (NFKC normalization, lowercasing, accent stripping).
-
Performing N-gram Hashing using deterministic functions ϕn,k to map canonicalized sequences to fixed memory indices in the shared memory table E.
-
Freezing the source Engram memory table EA and training only a lightweight target-side reader W(B) for extraction.
The Target-Side Reader Design: The Determinant of Performance
The performance ceiling is determined by the reader's design, which is treated as a first-order factor.
Key architectural choices include:
-
Memory Placement: Varying the layers (e.g., single-layer vs. dual-layer injection) at which memory is injected into the target backbone.
-
Retrieval Capacity: Increasing the number of reader branches (R), such as moving from R=1 to R=4, which
increases target-side alignment capacity without increasing memory size.
-
Reader Structure: Utilizing a
shared-value reader family
where different branches operate on the same retrieved content but differ in their key-gating pathways.
Evidence Chain: Reader Design Drives Gains
The evidence chain demonstrates that reader design is critical, showing that consistent gains across diverse model families and settings
are observed. Specifically:
)&Reader Placement:
**) Dual-layer injection (e.g., layers 2 and 10) increases average QA accuracy from 34.2 to 37.5 on LLaMA models, suggesting multiple integration opportunities are more beneficial than a single late-layer interface. **
) Reader Capacity:
**) Increasing branches from R=1 to R=4 raises average QA accuracy from 37.5 to 38.5, demonstrating that multi-branch gating provides complementary key-gating pathways over the same shared value representation.
**
Downstream Utility and Limitations
The utility of transferred memory is task-dependent:
-
Downstream Gains: Transferred memory shows
selective rather than universal
benefits; it improves factual and evidence-related tasks like RTE, SciQ, and BoolQ, but haslittle effect on broader reading comprehension.
-
TruthfulQA Boundary: Performance decreases on TruthfulQA across target scales (ranging from −0.2 to −0.8 accuracy points), suggesting that transferred memory is
most beneficial when the task rewards access to factual or evidence-related information, but it may be neutral or mildly harmful when success depends on calibration.
-
Efficiency: Frozen-memory transfer is
target-data efficient,
reaching strong intrinsic performance withsubstantially fewer target-side updates than learning a new memory from scratch,
though asymptotic ceilings are not necessarily higher than scratch training.
Conclusion: The Evaluation Framework
The paper concludes that cross-model frozen-memory extraction turns external memory portability into an evaluation framework. A memory artifact should be judged by whether a new backbone can address it, align to it, and extract useful signal without retraining the stored table.
The main takeaway is that the downstream gains cannot be explained by parameter count alone: successful transfer depends on meaningful memory addressing and a sufficiently expressive reader interface.
Furthermore, the cost analysis shows that the one-time provider cost is amortized after approximately nine adapted consumers.
Key Takeaways:
) Portability depends not only on the frozen table but also on the target-side reader.
**) Reader design is a "first-order determinant of transfer quality.
Improvements for AI systems
Based on the provided scientific paper, here are specific, actionable improvements to AI systems derived from the cross-model frozen-memory transfer (Engram) framework:
The core improvement is shifting knowledge utilization from brittle, model-specific weights to a reusable, addressable external knowledge artifact. This enables knowledge portability
across different model backbones and architectures.
Here are the specific improvements and capabilities:
-
Cross-Model Knowledge Transfer without Fine-Tuning (Zero Target Training):
-
Mechanism: Instead of retraining a target model on a new corpus to learn facts, you can transfer
frozen
memory from a source model (e.g., LLaMA) to a target model (e.g., Mistral). The only training required for the target is fitting a small, lightweightreader.
-
Capability: This allows deploying knowledge learned by Model A into Model B instantly, bypassing the massive computational cost and data requirements of full fine-tuning for every new task or model variant.
-
Robustness to Tokenizer Mismatches:
-
Mechanism: The system uses a
tokenizer-agnostic canonicalization pipeline
(NFKC normalization, lowercasing, accent stripping) before hashing memory addresses. This ensures that the same piece of text maps to the same memory slot regardless of whether Model A and Model B use different tokenizers (e.g., GPT vs. LLaMA). -
Capability: The transferred knowledge remains accessible across diverse model families and tokenizer boundaries, solving a major bottleneck in cross-model reuse that current RAG or adapter methods struggle with when architectures diverge significantly.
-
Optimized Knowledge Extraction via Reader Architecture Design:
-
Mechanism: The target system can be designed with a multi-layer, multi-branch reader (e.g., dual-layer injection at layers 2 and 10 with R=4 branches). This reader is optimized to intelligently
gate
and integrate the retrieved memory vector into the target model's hidden state. -
Capability: The AI system gains significantly higher accuracy (up to a 19.9% improvement on QA tasks) by learning the optimal way to interpret and route external knowledge, effectively creating a highly specialized
knowledge extraction module
tailored for the specific target backbone's residual stream. -
Task-Specific Knowledge Specialization:
-
Mechanism: The system can be tuned for specific downstream goals (e.g., high factual retrieval vs. resistance to misinformation) by controlling the training corpus used during Phase 2 reader adaptation (e.g., training on a QA-heavy corpus like Nemo HQ-DQA).
-
Capability: This allows the transferred memory to exhibit superior performance on tasks where factual evidence is key (like NQ, WebQA, and SciQ) while maintaining neutral or even slightly negative performance on tasks requiring calibrated truthfulness (like TruthfulQA), providing a fine-grained control over knowledge utility.
-
Data-Efficient Initialization:
-
Mechanism: The transferred memory provides a superior
structured initialization
compared to training a new memory from scratch, reaching competitive performance with significantly fewer target-side adaptation tokens (e.g., 5M tokens vs. 34M parameters). -
Capability: This drastically reduces the required data budget for adapting a new model to utilize existing knowledge, making knowledge reuse highly data-efficient and accelerating model deployment cycles.
-
Dynamic Knowledge Management:
-
Mechanism: The external memory is addressable and explicit; entries can be inspected, removed, or replaced without ever touching the backbone weights.
-
Capability: This enables modular knowledge updates—you can swap out specific facts or knowledge modules to improve performance without requiring a full model retraining cycle, offering unprecedented auditability and flexibility in managing learned information.
Sources
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Titans: Learning to Memorize at Test Time
- Pythia: A Suite for Analyzing Large Language Models Across Training and Scaling
- Improving language models by retrieving from trillions of tokens
- Transferring Linear Features Across Language Models With Model Stitching
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
- LoRA-X: Bridging Foundation Models with Training-Free Cross-Model Adaptation
- LoRA: Low-Rank Adaptation of Large Language Models
- Mistral 7B
- Generalization through Memorization: Nearest Neighbor Language Models
- Similarity of Neural Network Representations Revisited
- Large Memory Layers with Product Keys
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Exploiting Similarities among Languages for Machine Translation
- Qwen3 Technical Report
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- MLP Memory: A Retriever-Pretrained Memory for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering