Posterior contraction rates in Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families
University of Pavia · University of Turin · Collegio Carlo Alberto, Turin
math.ST, stat.ML, stat.TH
Submitted: 2026-08-11
Updated: 2026-10-05
Comments: 55 pages
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: The paper studies posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families.
Terminology
Summary
The paper studies posterior contraction in positive-order Sobolev norms and Bayesian derivative estimation for infinite-dimensional exponential families. The authors embed the natural parameter in a Hilbert scale and model it via a standard Gaussian series prior expanded in the eigenbasis generating the scale. Under a two-sided link condition on the Fisher information and suitable local regularity assumptions, they show that smoothness-matching priors achieve minimax-optimal posterior contraction rates in any Hilbert scale norm up to the regularity of the ground truth.
The analysis builds on the novel approach to posterior contraction based on the Wasserstein distance recently introduced by Dolera et al. (2024b). It combines refined Laplace-type estimates for infinite-dimensional integrals associated to the posterior kernels with a mixed-geometry estimate controlling their stability under fluctuations in the data, itself resting on a tailored Poincaré inequality for posterior distributions conditioned on neighbourhoods of the truth.
The main result, Theorem 3.5, shows that if the ground truth is in Hβ, the posterior distribution resulting from an α-regular Gaussian series prior contracts in Hs-norm at rate n−(α∧β−s)/(2α+d), for all 0 ≤ s < α ∧ β. Here, d ∈ N is an effective dimension index encoded by the asymptotic growth of the eigenvalues of the scale generator. For smoothness-matching priors (i.e., α = β), the obtained Hs-rate is thus equal to n−(β−s)/(2β+d) for all 0 ≤ s < β, coinciding with the usual minimax rate of estimation in Sobolev norms of order s for β-smooth functions defined on d-dimensional domains (Stone, 1982).
The proofs build on the novel approach to posterior contraction introduced by Dolera et al. (2024b), which bypasses testing arguments by directly controlling the expected Wasserstein distance between the posterior distribution and the Dirac mass at the ground truth, with transport cost induced by the metric of interest. This distance is upper bounded by the sum of a deterministic concentration term, governed by the local interaction between the likelihood and the prior near the true parameter, and a stochastic stability term that reflects the sensitivity of the posterior to fluctuations in the data. The deterministic term is controlled through Laplace-type integral bounds (Wong, 2001), refining the analysis of Dolera et al. (2024b) by replacing their oracle requirement that the prior covariance operator and the Fisher information at the ground truth be simultaneously diagonalisable with a two-sided link condition comparing the latter with a fixed power of the scale generator. For the stochastic component, the authors prove a mixed-geometry stability estimate for the posterior kernels arising in infinite-dimensional exponential families, which rests on a Poincaré inequality for the posterior distribution conditioned on a neighbourhood of the truth, established through Galerkin approximation and Brascamp–Lieb-type arguments (Bakry et al., 2008). The proof carefully decouples the strong geometry of the loss from the weaker geometry in which the sufficient statistic concentrates, thereby avoiding the algebraic loss in the rate incurred by the single-geometry strategy of Dolera et al. (2024b).
The general theory is applied to three concrete statistical models: density estimation with a logistic parametrisation, Poisson intensity estimation with an exponential link, and the Gaussian white-noise model. Under smoothness-matching Gaussian series priors, the authors obtain minimax contraction rates in Sobolev norms of every order (up to the regularity of the ground truth) for all three settings. Consequently, differentiated push-forward posteriors optimally recover the corresponding derivatives of the natural parameter, and, through the employed smooth parametrisation, also the derivatives of the target p.d.f. or intensity function.
For density estimation, the obtained results extend the B-spline analysis of Shen and Ghosal (2017) to Gaussian series priors expanded in standard bases generating the Sobolev scale, such as the Fourier and wavelet bases. In particular, the case s = 1 yields posterior contraction towards the true score function at the minimax rate n−(β−1)/(2β+d), or equivalently at the squared rate n−2(β−1)/(2β+d) relative to Fisher divergence (Wibisono et al., 2024). To the best of the authors' knowledge, the application to Poisson processes provides the first optimal contraction rates for derivatives of a nonparametric intensity function.
Improvements for AI systems
Improvements to AI Systems:
-
Uncertainty-Aware Bayesian Nonparametric Regression with Derivative Estimation: AI systems can be enhanced to perform nonparametric regression or classification while providing calibrated posterior uncertainty (credible intervals) not only for the function itself but also for its derivatives (e.g., velocity, acceleration, gradients) up to any order less than the function’s smoothness. This is directly useful for robotic control, time-series forecasting, and physics-informed machine learning, where derivative accuracy is critical.
-
Optimal-Rate Gaussian Process (GP) Priors for Sobolev-Smooth Targets: The theory enables AI systems to automatically select GP prior regularity (α) to match the unknown smoothness (β) of the target function, achieving minimax-optimal posterior contraction in any Sobolev norm. This improves GP-based Bayesian optimization and active learning by ensuring the posterior mean and uncertainty estimates converge at the fastest possible rate, avoiding over-smoothing or under-smoothing.
-
Robust Bayesian Inference for Infinite-Dimensional Exponential Families: AI systems can be built to handle models like logistic density estimation, Poisson intensity estimation, or Gaussian white-noise models with theoretically guaranteed posterior contraction. This is applicable to spatio-temporal point process models (e.g., neural spike trains, earthquake forecasting) and density estimation in high-dimensional feature spaces, where the system can reliably estimate intensity functions and their spatial/temporal derivatives.
-
Wasserstein-Based Posterior Concentration for Stability in Nonparametric Learning: The paper’s method of bounding posterior contraction via Wasserstein distance (instead of testing arguments) can be integrated into AI systems for continual learning or online Bayesian updating. The improved system would have provable stability guarantees under data perturbations, ensuring that posterior updates do not collapse or diverge, even when the likelihood and prior are not simultaneously diagonalizable—a common issue in deep kernel learning or neural network Gaussian processes.
-
Derivative-Aware Model Selection and Hyperparameter Tuning: By providing contraction rates for all Sobolev norms (including s=1 for score functions), AI systems can use posterior contraction of derivatives as a model selection criterion. For example, in generative modeling (e.g., score-based diffusion models), the system can tune the prior smoothness to achieve minimax-optimal estimation of the score function, improving sample quality and training stability.
-
Mixed-Geometry Stability for High-Dimensional Bayesian Deep Learning: The mixed-geometry stability estimate (decoupling strong loss geometry from weaker statistic concentration) can inspire new AI architectures for Bayesian neural networks. An improved system could use a Poincaré-inequality-based regularizer to control posterior fluctuations, leading to better generalization bounds and more reliable uncertainty quantification in high-dimensional parameter spaces.
-
Optimal Derivative Recovery for Physical Systems: For AI systems modeling physical phenomena (e.g., fluid dynamics, elasticity), the theory guarantees that push-forward posteriors of derivatives (e.g., stress, strain, velocity gradients) achieve minimax-optimal rates. This enables AI-based digital twins to provide accurate derivative fields from noisy observations, which is essential for design optimization and safety-critical monitoring.
Sources
- On strong posterior contraction rates for Besov-Laplace priors in the white noise model
- Adaptive Supremum Norm Posterior Contraction: Wavelet Spike-and-Slab and Anisotropic Besov Spaces
Related papers
- Conformal Prediction for Dyadic Regression Under Complex Missingness
- Bentkus-type asymptotic e-values
- High-Dimensional Asymptotics of Differentially Private PCA
- KL Convergence Guarantees for Score diffusion models under minimal data assumptions
- Geometric bias in eigenspace perturbation under random heterogeneous noise
- On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models