Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
cs.LG, cs.AI, cs.CL
Submitted: 2026-04-29
Updated: 2026-09-02
Comments: Also see arXiv:2505.21777 for a related work
Journal ref: Pham B. et al. Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data. In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing
License: http://creativecommons.org/licenses/by/4.0/
The gist: When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete
Terminology
Abstract
When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fundamentally behave as Associative Memories (AMs) with emergent creative capabilities. The core idea of an AM is to reliably recover stored data points as memories by establishing distinct basins of attraction around them. Historically, models like Hopfield networks use an explicit energy function to guarantee these stable attractors. We broaden this perspective by leveraging the observation that energy is not strictly necessary, as basins of attraction can also be formed via conditional likelihood maximization. By evaluating token recovery of training and test examples, we identify in UDDMs a sharp memorization-to-generalization transition governed by the size of the training dataset: as it increases, basins around training examples shrink and basins around unseen test examples expand, until both later converge to the same level. Crucially, we can detect this transition using only the conditional entropy of predicted token sequences: memorization is characterized by vanishing conditional entropy, while in the generalization regime the conditional entropy of most tokens remains finite. Thus, conditional entropy offers a practical probe for the memorization-to-generalization transition in deployed models.
Sources
- A Reproducible Extraction of Training Images from Diffusion Models
- Generalization in diffusion models arises from geometry-adaptive harmonic representations
- An analytic theory of creativity in convolutional diffusion models
- Losing dimensions: Geometric memorization in generative diffusion
- Understanding and Mitigating Memorization in Generative Models via Sharpness of Probability Landscapes
- Memorization to Generalization: Emergence of Diffusion Models from Associative Memory
- Memorization and Generalization in Generative Diffusion under the Manifold Hypothesis
- Generative Flows on Discrete State-Spaces: Enabling Multimodal Flows with Applications to Protein Co-Design
- Modern Methods in Associative Memory
- The Rules-and-Facts Model for Simultaneous Generalization and Memorization in Neural Networks
- Hierarchical Associative Memory
- Memory in Plain Sight: Surveying the Uncanny Resemblances of Associative Memories and Diffusion Models
- Deriving Neural Scaling Laws from the statistics of natural language
- One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling
- Simple and Effective Masked Diffusion Language Models
- Tractability from overparametrization: The example of the negative perceptron
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks