MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup
cs.CL, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Github: https://github.com/jojo23333/Mixutre-Of-Memory-Embedding
Code: https://github.com/jojo23333/Mixutre-Of-Memory-Embedding
License: http://creativecommons.org/licenses/by/4.0/
The gist: Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the
Terminology
Abstract
Scaling large language models efficiently has motivated sparse capacity mechanisms such as Mixture-of-Experts and, more recently, conditional memory: token-indexed embedding tables that augment the backbone with cheap parametric lookups. Existing memory-embedding methods retrieve via a deterministic function of the surface form, which collapses different contextual senses of the same token (e.g., python the language vs. the animal) into a single fixed entry. We introduce Mixture of Memory Embeddings (MoME), a context-aware memory mechanism that replaces each token's single memory row with a mixture of M slots and uses a learned gate over the hidden state to choose which slots to read at each position. In controlled pretraining experiments across nanochat, Llama-3/MobileLLM, and Qwen3 backbones, MoME improves over Value Embedding, Bigram, and STEM baselines in iso-parameter and iso-training-FLOP settings, shows a more promising memory-size scaling trend at sub-billion scale, and remains efficient in training and inference. Qualitative routing analyses on polysemous tokens further suggest that the learned mixture exhibits a degree of semantic interpretability, dispatching the same surface token to distinct memory slots under different senses.
Sources
- Memory Layers at Scale
- Improving language models by retrieving from trillions of tokens
- Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Knowledge Neurons in Pretrained Transformers
- DeepSeek-V3 Technical Report
- GLaM: Efficient Scaling of Language Models with Mixture-of-Experts
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Gemma 3 Technical Report
- Transformer Feed-Forward Layers Are Key-Value Memories
- The Llama 3 Herd of Models
- REALM: Retrieval-Augmented Language Model Pre-Training
- Mixture of A Million Experts
- Training Compute-Optimal Large Language Models
- Ultra-Sparse Memory Network
- Mixtral of Experts
- Scaling Laws for Neural Language Models
- Large Memory Layers with Product Keys
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering