From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models

arXiv:2609.39648 · cs.LG, cs.AI, cs.CV · Submitted 2026-09-30 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "From Modes to Memories".

Jane: Diffusion models can be interpreted as a family of deterministic dynamical systems indexed by noise scale, allowing researchers to characterize memorization through geometric dynamics.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So folks, we've been diving into this paper from arXiv called "From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models." The big idea here is framing diffusion models not just as random noise transformers, but as a family of deterministic systems depending on the noise scale. Jane, can you give us the simple breakdown?

Jane: Absolutely, Tom. Essentially, this paper proposes looking at a diffusion model by fixing the noise level and treating the denoiser like a map that keeps feeding its own output back into itself to see what happens over time. They claim that for an exact denoiser, fixed points are linked to critical points of the smoothed data density, and attractors correspond to modes; as you increase that noise scale, these sample-level modes start merging into broader ones.

Lu: That perspective shift is really interesting because it moves away from just looking at the final output and starts analyzing the underlying dynamics of how those outputs are generated across different scales. It opens up a geometric way to view things like memorization.

Meng: From an engineering standpoint, if we can characterize these modes and attractors, it tells us something about the stability of the learned representations at different levels of noise. But how does this map onto what we see in a real trained network?

Lalam: I think this geometric view is powerful because it gives structure to phenomena we usually only observe statistically, like why some images are memorized and others aren't. It’s providing a concrete mathematical language for something fuzzy.

Tom: Exactly, Jane, so the core claim is that examples that get extra probability mass because of duplication or outliers should behave differently when we apply stronger smoothing; they should stay distinguishable under larger noise levels. This leads directly to their key measure: the critical scale, sigma c(x).

Jane: That critical scale is what they define as the largest noise level where an example is still retained by those fixed-scale dynamics; it's this persistence measure. They test this by checking if the orbit stays close to the original point x, specifically looking at when Mk sigma(x) - x rx for every iteration k up to some limit K.

Lu: The paper shows that sigma c(x) captures memorization arising from duplication, overfitting, and outliers in the data distribution. It’s not just a single concept; it ties these distinct memorization mechanisms together through this persistence measure.

Meng: I see how that connects to the practical side of things; if sigma c tells us something about persistence, we can use it to test if our models are truly generalizing or just memorizing specific training points. But what about conditional models, like image-caption pairs?

Paper summary: Lalam: That's where things get really insightful; in conditional settings, the paper introduces the "caption gap," defined by sigma c(x; c) = sigma c(x; c) - sigma c(x;).

Tom: So if that caption gap is positive, it means the caption actually increases the model's mass on that specific example, which directly characterizes memorization arising from duplication or overfitting in those conditional contexts. That's a really sharp distinction.

Jane: It gives us a way to isolate whether the memorization is happening because of just the image itself or if the caption is actively contributing to that retention. This moves beyond just looking at raw data points and into understanding how different data modalities interact within the model’s learned structure.

Lu: The scaling laws they establish, like doubling the distance to a neighbor doubling sigma c while doubling copies only raises it by a few percent, offer some very specific predictions about how sensitive this measure is to different types of memorization. It suggests that spatial arrangement matters more than sheer quantity when it comes to duplication effects.

Meng: That’s useful for setting up controlled experiments; if we know that distance has a more pronounced effect on sigma c than the number of copies, we can focus our efforts on how close examples are clustered rather than just adding massive amounts of redundant data.

Lalam: And if we look at the local gaps, sigma c(u), we can spatially map where the memorized content is localized within an image, identifying regions that are reproduced consistently across different noise levels. That’s a really tangible way to visualize the 'where' of memorization.

Tom: It seems like we've got a solid handle on the mechanism now; it’s not just that something is memorized, but *how* it persists under different noise scales, and how context like captions affects that persistence. This paper really gives us the tools to probe the inner workings of these massive diffusion models.

Jane: It does provide a theoretical bound as well; they compare their critical scale derived from the exact denoiser, sigma empc(x), with what we observe in trained networks using the retention coefficient, kappa(x) = sigma c(x) / sigma empc(x).

Lu: That comparison is key because it shows that this theoretical scale provides an upper limit on what any trained model can achieve, meaning the mechanism transfers to them, but the exact denoiser's scale sets a benchmark for all of them.

Paper summary: Meng: From a practical perspective, that means we have a way to mathematically bound how much memorization we can expect even in highly optimized networks. It helps us understand the inherent limits of generalization versus rote memorization.

Lalam: If this framework is applied broadly, it could lead to better ways of auditing model behavior, helping us understand why some creative outputs are robust and others are just brittle copies. It gives us a vocabulary for discussing the cultural impact of these systems.

Tom: So, to wrap up this part, the authors have given us a powerful lens to examine diffusion models as dynamical systems indexed by noise scale, using sigma c to quantify persistence and conditional gaps to analyze caption influence.

Jane: And the conclusion is that this critical scale offers an effective global measure of memorization at a given scale, and computing it locally lets us pinpoint exactly where in an image the memorized content appears.

Lu: The implication here is that we can move from simply observing outputs to understanding the geometric reasons behind their stability or instability under model operation. It suggests a much richer structure beneath the stochastic noise.

Meng: For deployment, this means we have a metric, sigma c, that isn't just an arbitrary number but one tied directly to data density and distance. That makes it useful for setting safety thresholds related to model fidelity.

Lalam: I think the cultural implication is that understanding this persistence allows us to build models with more intentionality, knowing exactly where the boundaries of memorization lie in generated content. It gives us a foundation for building more reliable and less arbitrary systems.

Tom: What we're hearing is that this paper gives us a rigorous mathematical tool to dissect the behavior of diffusion models, moving beyond simple correlation to geometric persistence.

Jane: And it connects the theoretical structure of these dynamical systems directly to observable phenomena like how examples are retained across noise levels.

Lu: It provides a framework for characterizing memorization through geometric dynamics, which is a really deep way to look at this type of AI architecture.

Meng: So we can use sigma c to guide how we train and fine-tune these models, aiming for better generalization instead of just hoping for the best.

Lalam: It’s exciting because it gives us a way to interpret the complex behavior of these large generative systems in terms of underlying data structure and dynamics.

Tom: We've covered the core concepts and the implications for understanding how these models learn, setting up a really solid foundation for what we're hearing next.

Conclusion: Tom: So, we’ve been digging into how diffusion models act like dynamical systems depending on noise scale, and now we’re at the conclusion of this paper, "From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models."

Jane: Exactly. This research boils down to understanding how an AI model remembers things by tracking examples through different levels of noise. It really moves beyond just looking at the final picture and gets into what’s happening beneath the surface during that generation process.

Lu: The authors do a lot of heavy lifting here, framing the whole diffusion process as a family of maps indexed by noise scale, which is a really deep way to look at how these models operate internally.

Meng: I'm curious, Tom, when you look at this conclusion about the critical scale sigma c, how does that translate into something we can actually measure in terms of model performance or safety?

Lalam: The most impactful vision I have here is that we can now characterize memorization through a geometric lens, allowing us to isolate precisely *where* and *why* an AI retains specific training data points.

Tom: That’s right, Lalam. And the authors show that this critical scale sigma c isn't just some abstract number; it’s a concrete measure of how long an example stays "sticky" under the model's own smoothing process.

Jane: It helps us see that memorization isn't one single thing, but rather a spectrum of phenomena like duplication or outliers, each with its own persistence characteristics defined by this scale.

Lu: I think it’s fascinating because they also provide a way to test if the mechanism learned in theory actually transfers over to the trained networks we use every day.

Meng: That transferability is crucial for practical application; if the theoretical dynamics hold up in practice, then we can have more confidence in how well we can tune these models for specific tasks.

Lalam: This kind of deep structural understanding could profoundly impact culture because it gives us a way to audit generative content for patterns of rote memorization versus true creative synthesis.

Tom: And that’s the big picture, folks. So, this paper really gives us the tools to map the hidden dynamics of diffusion models using these scale-space concepts.

Jane: It’s a powerful framework for understanding how these complex AI systems build their knowledge base and what keeps those pieces locked in place across different noise levels.

Lu: We have a lot more to unpack regarding how this geometric interpretation can be used to design new ways of training generative models for better generalization.

Cristina López Amado, Marco Fumero, Francesco Locatello

Institute of Science and Technology Austria (ISTA)

cs.LG, cs.AI, cs.CV

Submitted: 2026-09-30

Updated: 2026-09-30

Code: https://github.com/huggingface/diffusers

Importance score: 83/100

The gist: Diffusion models can be interpreted as a family of deterministic dynamical systems indexed by noise scale, allowing researchers to characterize memorization through geometric dynamics.

Key concepts

Critical Scale ($\sigma_c(x)$)
This is the largest noise level at which an example is retained by the fixed-scale denoiser dynamics. It serves as a measure of memorization, indicating how robust an example is to being lost during the denoising process. It captures phenomena like overfitting or outliers in data.
Scale Space of Attractors
As the noise scale ($\sigma$) increases, the set of attractors (modes) in the denoiser map changes. This concept describes how different modes merge into coarser ones as noise becomes more significant, showing how the model's learned features evolve across different levels of noise.
Caption Gap ($\Delta \log \sigma_c(x; c)$)
In conditional models (like image-caption pairs), this gap measures the difference in critical scales between an example with a caption and one without. A positive gap means the caption helps the model retain more mass on that specific example, characterizing memorization arising from caption influence.

Terminology

Summary

Diffusion models can be interpreted as a family of deterministic dynamical systems indexed by noise scale, allowing researchers to characterize memorization through geometric dynamics. The critical scale, or persistence measure, quantifies how long an example remains an attractor under iteration of the fixed-scale denoiser, thereby providing a framework to distinguish between different types of memorization.

How it works

The paper proposes viewing a diffusion model as a family of denoisers, denoted as the set of maps Mσ, where each map Mσ is defined by fixing the noise level σ and iterating the denoiser on its own output: xk+1 = Dσ(xk). For an exact denoiser, this map is described by Proposition 2 as a Gaussian mean shift on the smoothed data density pσ(x), with fixed points corresponding to critical points of pσ and attractors corresponding to its modes. Varying the noise scale σ traces a scale space of the attractors, where modes merge into progressively coarser ones as σ increases.

The Critical Scale as a Memorization Measure

The critical scale, denoted as σc(x), is defined as the largest noise level at which an example is retained by the fixed-scale dynamics. This measure captures memorization arising from several phenomena:

  1. Duplication or overfitting of training data.

  2. Outliers in the data distribution.

The persistence of an example under its own dynamics is tested by checking if the orbit remains close to it, defined as: the probe x is retained at scale σ if Mkσ(x) − x ≤ rx for every k ≤ K.

Quantifying Memorization Mechanisms

The critical scale provides interpretable measures of memorization. The paper shows that σc responds to the probability mass carried by an example and its distance from the rest of the data. Specifically, Theorem 5 establishes a scaling law: Doubling the distance to the neighbour therefore doubles σc, whereas doubling the copies raises it by a few percent. This indicates that duplication increases σc only slightly while increasing distance has a more pronounced effect.

Conditional Models and Caption Dependence

In conditional models (like image–caption pairs), memorization becomes a property of the pair. The paper introduces the caption gap, defined as ∆ log σc(x; c) = log σc(x; c) − log σc(x ∅). Corollary 7 shows that this gap is positive if and only if the caption raises the model’s mass on that example. This allows for the characterization of memorization arising from duplication, overfitting, and outliers in conditional settings.

Detecting Memorization in Controlled Experiments

Controlled experiments were conducted to transfer the theoretical mechanism to trained networks. The paper tested interventions such as:

  1. Duplication: Repeating an example m times in the training set.

  2. Rarity: Adding atypical examples with little training neighbor data but low probability mass.

  3. Overfitting: Reducing the training set size at a fixed model capacity and budget to move from generalization to memorization, where σc tracks this transition and detects memorization when sampling no longer finds copies.

Interpretation of Local Gaps

The paper demonstrates that the critical scale can be localized spatially. By computing the local gap ∆ log σc(u), researchers can identify where in the image the memorized content localizes, as large values of this gap align with regions reproduced consistently across generations. This provides a spatial measure of where memorized content appears within an image.

Comparison to Trained Networks

The paper compares the critical scale derived from the exact denoiser (Mempσ) with that observed in trained networks. The retention coefficient, κ(x) = σc(x) / σempc(x), measures the fraction of that solution’s retention the network reaches at x. This comparison shows that while the mechanism transfers to trained networks, the exact denoiser's critical scale provides a theoretical bound: the memorizing solution therefore bounds every trained model we measure.

Conclusion

The paper concludes that diffusion models can be interpreted as a family of dynamical systems indexed by noise level, and the critical scale σc measures the persistence of an example as an attractor. On Stable Diffusion, σc detects fully and partially memorized image–caption pairs, offering interpretations by comparing conditional and unconditional critical scales to isolate caption-image interaction. Furthermore, local variants of σc can identify which regions within an image are memorized. The theory holds for the exact denoiser, and the mechanism transfers to trained networks. The critical scale is an effective global measure of memorization at scale, and computing it per spatial location improves detection.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed this paper, From Modes to Memories: Characterizing the Scale-Space Dynamics of Diffusion Models. The core contribution is shifting the perspective from diffusion models as purely stochastic processes to a family of deterministic dynamical systems indexed by noise scale. This framework introduces the concept of the critical scale, σc, as a rigorous measure of memorization.

Here are specific improvements that can be made to AI systems based on this research:


)

Improving Model Interpretability and Debugging

The paper establishes a clear mathematical framework for how training data is retained or memorized within a diffusion model's dynamics (attractors). This allows for the development of diagnostic tools:

  1. [Critical Scale Measurement] Implement a quantitative metric, based on the critical scale σc, to automatically score generated images against known memorized examples.

  2. [Memorization Localization] Utilize local variants of σc (the local critical-scale gap, ∆ log σc(u)) to spatially map which specific regions of an image are memorized (e.g., the curtain pattern vs. the bedsheet pattern).

  3. [Attractor Persistence Analysis] Track how long a generated sample remains an attractor under different noise levels (σ). This can quantify the robustness of learned features versus noise, helping distinguish between genuine generalization and transient retention.

)

Enhancing Training Data Curation and Quality Control

The paper provides mechanisms to identify specific failure modes during training:

  1. [Duplication Detection] Use σc to automatically detect training examples that are exact copies or near-duplicates (multiplicity m > 1), allowing for targeted filtering of redundant data before retraining.

  2. [Outlier Identification] Apply the rarity detection threshold derived from Theorem 5 to flag atypical training samples (outliers) that might be disproportionately influencing the model's dynamics, enabling better data balancing during training.

  3. [Overfitting Monitoring] Monitor the shift in σc as a function of training set size (N). A rapid decrease in σc signals the transition from generalization to memorization, providing an objective stopping criterion for training schedules.

)

Refining Conditional Generation and Captioning

The research extends to conditional models (image-caption pairs), offering insights into multimodal memorization:

  1. [Caption Impact Assessment] Calculate the caption gap (∆ log σc) between the unconditional model and a conditional model. A large positive gap indicates that the caption substantially increases retention, allowing developers to understand which captions are responsible for locking in specific visual features.

  2. [Spatial Caption Localization] Use local ∆ log σc maps on image-caption pairs to determine if memorization is driven by global caption reliance or localized spatial content dependence within the image (e.g., whether a caption affects the curtain region more than the bedsheet).

)

Developing Robust Model Auditing and Safety Systems

The framework directly addresses membership inference and data recovery concerns:

  1. [Membership Inference Detection] Use σc to distinguish between training images that are strongly retained (memorized, high σc) and those that are transient (transient, low σc), offering a more sensitive measure than traditional membership inference attacks.

  2. [Unconditional Retention Evidence] Employ the analysis of the signed caption gap to identify images that retain their content even when conditioned on a mismatched or empty caption (strongly negative gap), providing evidence for memorization occurring in the unconditional branch, which is crucial for safety auditing.

)

Specific System Capabilities Enabled by These Improvements:

  1. [High-Fidelity Content Verification] A system can ingest a generated image and report: This image exhibits high memorization in the 'sky' region (high local ∆ log σc), likely due to a specific caption.

  2. [Automated Data Deduplication Pipeline] A pipeline that analyzes new training data against the existing set, flagging any new images with multiplicity m > 1 for automated de-duplication or quality assessment.

  3. [Adversarial Robustness Tuning] By monitoring σc during fine-tuning, a system can dynamically adjust regularization (like dropout or weight decay) to prevent the critical scale from collapsing prematurely due to overfitting.

Sources

Related papers