Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions

arXiv:2409.18804 · stat.ML, cs.LG, math.ST, stat.TH · Submitted 2026-08-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions".

Jane: The paper was written by Iskander Azangulov, Judith Rousseau and George Deligiannidis from University of Oxford and Universite Paris-Dauphine and PSL University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s been making waves in the machine learning world: "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions."

Jane: And Tom, I have to say, this one got me excited right from the title. We’re talking about diffusion models—those are the engines behind image generators—and the question of why they work so well when the data lives on a low-dimensional structure hidden inside a huge ambient space.

Tom: Exactly. The paper’s from a team at Oxford and Paris-Dauphine, and they’re tackling something that’s been bugging theorists for a while. You’ve got data that’s nominally in, say, a thousand dimensions, but it actually sits on a much smaller manifold. The manifold hypothesis says that’s how real-world data behaves.

Jane: And the big question they ask is: can we prove that diffusion models adapt to that hidden low-dimensional structure? Not just in practice, but with actual mathematical guarantees that don’t blow up with the ambient dimension.

Tom: Right, and that’s the headline. Previous work had shown convergence rates that depended heavily on that big ambient dimension, which made the theory look bad compared to how well these models work in practice. This paper closes that gap.

Jane: So instead of saying “the error grows with the number of dimensions,” they show the error basically depends only on the intrinsic dimension of the data manifold. That’s a huge deal for understanding why these models are so successful.

Tom: And it’s not just a theoretical curiosity. If you’re generating images, the ambient dimension is millions of pixels, but the meaningful structure is way smaller. This result gives a rigorous reason why diffusion models can handle that.

Jane: Let’s hold that thought, because the next segment is going to dig into what they actually proved and how they got there. Stick around.

Summary: Tom: Welcome back. We’re still on "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions." Jane, you were just about to tell us what the paper actually delivers.

Jane: Right. So the paper gives two main results. First, they show that a neural network can learn the score function—that’s the mathematical object diffusion models use to reverse the noise—with an error that’s independent of the ambient dimension. The rate they get is essentially the optimal one for the intrinsic dimension.

Tom: And the second result builds on that. They show the whole sampling procedure, the discretized backward process, achieves the same near-optimal rate in the Wasserstein distance. That’s a measure of how close the generated samples are to the true data distribution.

Jane: And the key phrase there is “independent of ambient dimension.” Previous work had rates that scaled with the ambient dimension to some power, which meant the theory predicted terrible performance in high dimensions. This paper fixes that.

Tom: How do they do it? I mean, that’s the part that’s really clever. They develop a new framework connecting diffusion models to the theory of extrema of Gaussian processes. That’s a classical area of probability theory.

Jane: And that connection lets them control how the added noise interacts with the manifold. They show that the noise is almost orthogonal to the manifold when the ambient dimension is large, so the score function only needs to be accurate in the low-dimensional directions.

Tom: So the noise is like wind blowing mostly sideways, and the model only needs to learn the small component that actually moves the data. That’s the intuition, right?

Jane: Exactly. And they also build a piecewise polynomial approximation of the manifold itself, using the training samples. That part is borrowed from earlier work on manifold estimation, but they adapt it to keep the dimension of each piece small.

Tom: Small enough that the neural network can learn each piece efficiently. That’s the practical trick that makes the whole thing work.

Jane: And the result is that the sampling complexity—how many steps you need—also doesn’t depend on the ambient dimension. Just the intrinsic one, up to log factors.

Tom: That’s a beautiful result. Next segment, let’s talk about what they actually improved compared to the existing literature.

Improvements: Tom: Back for more on "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions." Jane, you mentioned this improves on earlier work. Can you be specific?

Jane: Sure. The most relevant prior result was from a paper by Tang and Yang, who also studied diffusion models under the manifold hypothesis. They got the right rate in terms of the sample size, but their error bound had a factor of the ambient dimension raised to a power like d/two plus alpha.

Tom: And that’s a problem because if the ambient dimension is, say, a million pixels, that factor is astronomically large. The theory would say the model should fail, but in practice it works great.

Jane: Exactly. This paper removes that factor entirely. The error depends only on the intrinsic dimension and the smoothness of the data distribution, plus logarithmic terms in the ambient dimension.

Tom: So the improvement is not just a constant factor—it’s a change in the scaling behavior. That’s the kind of improvement that changes the qualitative prediction of the theory.

Jane: And they also improve the architecture. The previous work used neural networks that operated directly in the ambient space, which meant the number of parameters had to scale with that dimension. Here, they construct local coordinate systems of dimension O(log n), so the networks only need to work in those low-dimensional spaces.

Tom: That’s a clever engineering trick embedded in the theory. Each local patch of the manifold is approximated by a polynomial surface that lives in a subspace spanned by nearby training points.

Jane: And because each patch is low-dimensional, the neural network can approximate the score function on that patch without ever needing to see the full ambient space.

Tom: So the improvement is both in the bound and in the architecture. That’s a rare combination—theory and practice moving together.

Jane: And the proof technique is also novel. They use concentration inequalities for Gaussian processes to control the noise, and they combine that with a careful analysis of the posterior distribution of the denoising target.

Tom: Next segment, let’s look at the first page of the paper itself and see how they frame the problem.

First Page: Tom: We’re back on "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions." Jane, let’s look at the opening of the paper. What’s the framing?

Jane: The abstract sets up the problem beautifully. It starts with the manifold hypothesis—the idea that high-dimensional data often lies on lower-dimensional manifolds. And it notes that while recent results have given insight into how diffusion models adapt to this, they don’t capture the empirical success.

Tom: And that’s the gap they’re filling. The introduction goes on to explain that diffusion models work by gradually adding Gaussian noise and then learning to reverse that process. The score function is the key object.

Jane: Right. And they point out that the previous best bound had that strong dependence on the ambient dimension, which meant the error would go to infinity as the ambient dimension grows with the sample size. That’s clearly not what happens in practice.

Tom: So the paper’s contribution is to show that the normalized score function can be learned at a rate independent of the ambient dimension, up to log terms. And that leads to the optimal Wasserstein rate.

Jane: The introduction also mentions that they build on the manifold estimation work of Aamari and Levrard, and the measure estimation work of Divol. So it’s a synthesis of geometric statistics and diffusion model theory.

Tom: And they’re explicit that their approach connects to the theory of Gaussian processes, which is a fresh perspective in this area.

Jane: One thing I appreciate is that they’re honest about the limitations. They have a technical assumption relating the manifold smoothness to the measure smoothness, and they note it might be relaxable.

Tom: So the first page sets up a clear problem, states the main results, and gives the roadmap. It’s a well-written introduction.

Jane: And it sets the stage for the technical sections, which we’ve been discussing. Let’s wrap up with our final thoughts in the next segment.

Conclusion: Tom: And we’re at the end of our discussion on "Convergence of Diffusion Models Under the Manifold Hypothesis in High-Dimensions." Jane, what’s the big takeaway for our listeners?

Jane: The big takeaway is that this paper provides a rigorous explanation for why diffusion models work so well on high-dimensional data that has low-dimensional structure. The error rates are independent of the ambient dimension, which matches the empirical success.

Tom: And they achieve this by combining manifold estimation with Gaussian process concentration bounds, plus a clever low-dimensional architecture for the neural networks.

Jane: The implications are significant. For practitioners, it means the theory now supports what they’ve been seeing in practice—that these models can handle high-dimensional data without the curse of dimensionality.

Tom: And for theorists, it opens up new directions. The authors mention several open questions, like relaxing the smoothness assumption and handling noisy observations.

Jane: There’s also the possibility of extending these techniques to other diffusion model formulations, like stochastic interpolants. The authors note that their approach could be adapted.

Tom: So this paper is not just a single result—it’s a framework that could influence future work in generative modeling theory.

Jane: And for the world, it brings us closer to understanding why these powerful tools work, which is important as they’re used in everything from image generation to drug discovery.

Tom: Well said, Jane. That’s our show for today. Thanks for joining us, and we’ll see you next time with another exciting paper from arXiv.

Jane: Goodbye, everyone, and keep exploring the math behind the machines.

Iskander Azangulov, Judith Rousseau, George Deligiannidis

University of Oxford · Universite Paris-Dauphine · PSL University

stat.ML, cs.LG, math.ST, stat.TH

Submitted: 2026-08-06

Updated: 2026-08-10

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 51/100

The gist: - For d ≥ 3: EY∼µ⊗n ∫∫ σt2∥ŝ(t, x) − s(t, x)∥2 p(t, x)dxdt ≤ n−2(α+1)/(2α+d) polylog n - Otherwise: (log n)γ · n−1 Corollary 3.2 states that with T = log D + log n =

Key concepts

Diffusion Models
These are generative models, described as engines behind image generators. They work by gradually adding Gaussian noise to data and then learning a process to reverse that noise, allowing for sample generation.
Manifold Hypothesis
This hypothesis suggests that real-world high-dimensional data (like images) actually resides on a much smaller, low-dimensional structure called a manifold within the larger ambient space.
Ambient Dimension
This refers to the total number of dimensions in the space where the data is nominally placed (e.g., millions of pixels for an image). The paper's breakthrough was showing results independent of this large number.
Intrinsic Dimension
This is the true, underlying low-dimensional structure that describes how the data actually varies, even if it lives in a much higher ambient dimension.

Terminology

Summary

Summary

The paper studies Denoising Diffusion Probabilistic Models (DDPM) under the manifold hypothesis, which postulates that high-dimensional data often lie on lower-dimensional manifolds within the ambient space. The authors prove that diffusion models achieve convergence rates independent of the ambient dimension in terms of score learning and sampling complexity.

Main Results

The paper's primary contribution is showing that the normalized score function σt s(t, x) can be learned by a neural network estimator with the (near) optimal convergence rate n−(α+1)/(2α+d) up to polylog term w.r.t. the score matching loss, implying an optimal (up to polylog) rate n−(α+1)/(2α+d) in the Wasserstein metric as long as log D = O(log n). The ambient dimension D only impacts the rate via a logarithmic term.

Theorem 3.1 states that for a measure µ satisfying Assumptions A–D supported on a d-dimensional β-smooth manifold M embedded into RD, with T ≍ (log n)γ n−2(α+1)/(2α+d) and T = O(log n), there exists an estimator ŝ(t, x) satisfying:

  • For d ≥ 3: EY∼µ⊗n ∫∫ σt2∥ŝ(t, x) − s(t, x)∥2 p(t, x)dxdt ≤ n−2(α+1)/(2α+d) polylog n

  • Otherwise: (log n)γ · n−1

Corollary 3.2 states that with T = log D + log n = O(log n), there is a discretization scheme consisting of O(n(2α+d)/(α+1) polylog n) discretization steps achieving EY∼µ⊗n W1(µ, µ̂) ≲ polylog n · n−(α+1)/(2α+d) for d ≥ 3, and n−1/2 otherwise.

Key Innovations

  1. High-probability bounds on the score function: The authors develop a novel framework connecting diffusion models to the theory of extrema of Gaussian Processes. Proposition 4.1 shows that for any δ > 0 and ε < r0 with probability 1 − δ for all y, y′ ∈ M:

⟨ZD, y − y′⟩ ≤ 4ε√d + ∥y − y′∥ + 6ε√(4d log 2ε−1 + 4 log+ Vol M + 2 log 2δ−1)

  1. Dimension reduction scheme: The authors modify the manifold estimator from [3] to construct a dimension-reduction scheme allowing the use of n samples to build an efficient approximation of a β-smooth manifold M with error n−β/d, by n polynomial surfaces each contained in easy-to-find sub-spaces of dimension O(log n).

  2. Localization of the posterior distribution: Theorem 4.4 shows that with high probability the posterior mass of X0 given Xt is concentrated on BM(X0, rt) with rt ≍ (σt/ct)√(d log+(ct/σt) + Clog + log n).

Estimator Construction

For d ≤ 2, the empirical score function is used directly. For d ≥ 3, the construction is blockwise in time, splitting [T, T] into dyadic blocks [Tk, Tk+1] with Tk+1/Tk = 2. The estimator uses Tweedie's formula s(t, x) = (ct e(t, x) − x)/σt2, where e(t, x) is the conditional expectation of X0 given Xt.

The estimator is obtained by minimizing the empirical risk over neural networks with architecture:

ϕ(t, x) = (ct/σt2) · (Σi ρ̃i(t, x)ϕwi(t, x̃i,t)(Gi + PHi ϕei(t, x̃i,t))) / (Σi ρ̃i(t, x)ϕwi(t, x̃i,t)) − x/σt2

where ρ̃i are localization functions, Gi are sample points, Hi are low-dimensional subspaces (dim Hi = O(log n)), and ϕei, ϕwi are neural networks of polylogarithmic size.

Discretization Scheme

The paper presents a randomized discretization mesh (Section 6.1) with properties ensuring the discretized denoising loss can be controlled by its continuous version. The discretization scheme (6.11) uses a self-consistency modification of the denoiser to ensure stability along the trajectory.

Theorem 6.1 proves that under Assumption E (continuous denoising loss bounded by ε2denoiser and high-probability tail bounds), there exists a discretization scheme requiring L = O(ε−2denoiser log T + T) steps to generate samples from µ̂ satisfying:

W1(µ, µ̂) ≲ √De−T + √(T log T−1)√(CW + log(ε−1denoiser · T)) + εdenoiser · log(T · T−1)√(log T−1)√(CW + log(ε−1denoiser · T))√((d log T−1 + CW)(log T−1 + T))

Alternative Approach

Section 7 presents an alternative kernel-based score function approximation using the estimator of [15], achieving the same rate n−(α+1)/(2α+d) log n, though this approach merely samples from an already-known distribution rather than learning it through empirical risk minimization.

Assumptions

  • Assumption A: M is a compact β-smooth manifold of dimension d ≥ 1 with reach τ > τmin

  • Assumption B: µ has α-smooth density p(y) with pmin ≤ p ≤ pmax

  • Assumption C: Bounds on density, volume, and regularity constants via Clog

  • Assumption D: If d ≥ 3, β ≥ d/(d−2)(α+1)

Improvements for AI systems

Based on the paper, here are specific improvements that can be made to AI systems, particularly generative models and score-based diffusion models:

Improvement: Replace the current neural network architecture for score estimation with the localized, low-dimensional construction described in Section 3.1.2. The key innovation is to:

  • First, estimate the unknown low-dimensional manifold using the polynomial surface approximation from Section 5.1 (solving optimization problem 5.1)

  • Then, project the score estimation problem onto the learned local tangent spaces (dimension O(log n)) rather than working in the full ambient space R D

  • Use the localized weighting functions ρi(t,x) that restrict attention to relevant manifold patches

What the improved system can do: Achieve score-matching loss convergence rates of n-2(α+1)/(2α+d) polylog(n) that are independent of the ambient dimension D, even when D >> n. Current systems degrade as D α+d/2 and fail completely when D ≳ n.

Abstract

Denoising Diffusion Probabilistic Models (DDPM) are powerful state-of-the-art methods used to generate synthetic data from high-dimensional data distributions and are widely used for image, audio, and video generation as well as many more applications in science and beyond. The manifold hypothesis states that high-dimensional data often lie on lower-dimensional manifolds within the ambient space, and is widely believed to hold in provided examples. While recent results have provided invaluable insight into how diffusion models adapt to the manifold hypothesis, they do not capture the great empirical success of these models, making this a very fruitful research direction. In this work, we study DDPMs under the manifold hypothesis and prove that they achieve rates independent of the ambient dimension in terms of score learning. In terms of sampling complexity, we obtain rates independent of the ambient dimension w.r.t. the Wasserstein distance. We do this by developing a new framework connecting diffusion models to the well-studied theory of extrema of Gaussian Processes.

Sources

Related papers