A Local Sinkhorn Framework for Conditional Distribution Reconstruction of Multidimensional Random Fields
Mingtao Xia, Qijing Shen
cs.LG
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper proposes a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields.
Terminology
Summary
This paper proposes a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields. The proposed approach utilizes the debiased Sinkhorn divergence to develop a differentiable and computationally efficient local distribution matching objective to train stochastic neural networks (SNNs). The paper establishes theoretical generalization error estimates for the local Sinkhorn divergence framework, which explicitly characterizes the trade-off between approximation bias and statistical efficiency controlled by the regularization parameter and reveals how the proposed local Sinkhorn divergence loss function can be efficiently applied to learning multidimensional random field models. The proposed framework provides a scalable alternative to exact local optimal transport for conditional distribution reconstruction, offering a practical compromise between geometric fidelity, statistical efficiency, and computational scalability for uncertainty quantification and probabilistic scientific machine learning. Through various numerical examples, the paper compares the proposed local Sinkhorn divergence framework with other loss functions to train SNNs and with other machine-learning-based uncertainty quantification frameworks, demonstrating that the proposed local Sinkhorn divergence framework achieves an effective balance between reconstruction accuracy and computational efficiency while maintaining good scalability for multidimensional stochastic systems.
The main contributions of this work are summarized as follows:
• We extend our recent local optimal transport framework for random field reconstruction from the exact Wasserstein distance to the debiased Sinkhorn divergence, leading to a computationally efficient, scalable, and fully differentiable local distribution matching method for training SNNs to reconstruct random fields.
• We provide a theoretical analysis of the proposed local Sinkhorn formulation by combining the neighborhood approximation error with the regularization and statistical convergence properties of empirical Sinkhorn divergence. The resulting generalization bounds reveal how the local Sinkhorn divergence may partially alleviate the curse of dimensionality through introducing an entropic regularization term.
• Through various numerical examples, we demonstrate that the proposed local Sinkhorn divergence provides an effective compromise between approximation accuracy and computational efficiency and outperforms several machine-learning-based uncertainty quantification benchmarks, making it particularly suitable for multidimensional uncertainty quantification problems.
The paper introduces two theorems providing generalization error bounds. Theorem 2.1 provides a generalization error bound on estimating how the empirical Sinkhorn divergence loss function, when utilizing the neighborhood technique, approximates the averaged squared W2 distance between the ground truth random field and the approximate model. Theorem 2.2 provides another generalization error bound based on the generalization error of the Sinkhorn divergence. The bounds suggest an important practical implication: when the data are highly heterogeneous, it is preferable to choose a relatively small regularization parameter epsilon, so that the local Sinkhorn loss remains close to the local squared W2 loss. On the other hand, when the noise in the target variable is homogeneous, a moderately large regularization parameter can be advantageous, so that the convergence rate as the number of training samples increases is improved. As a result, by tuning epsilon, the local Sinkhorn loss provides a flexible trade-off between statistical accuracy and computational robustness, and may partially alleviate the curse of dimensionality compared with the exact local squared W2 loss.
Numerical experiments are carried out on three examples: one-dimensional conditional distribution reconstruction, stochastic Darcy flow, and stochastic FitzHugh–Nagumo (FHN) systems. In Example 1, the proposed local Sinkhorn loss achieves the smallest overall errors for both the conditional mean and conditional variance compared to local MSE, local MAE, local Energy Distance, local MMD, and local squared W2. In Example 2, training the SNN with the local Sinkhorn divergence achieves the smallest errors in the predicted mean and variance compared to Heteroscedastic Gaussian Regression, Mixture Density Network, Conditional VAE, Conditional Normalizing Flow, SNN + Local W2, and SNN + MMD, while also improving computational efficiency compared to the local squared W2 approach. In Example 3, using the Sinkhorn divergence is more computationally efficient, and the learned drift and diffusion functions are more accurate compared to those learned when using the temporally decoupled local squared W2 as the loss function.
Improvements for AI systems
Improvements to AI Systems:
-
Adaptive Regularization for Heterogeneous Data: Implement a dynamic tuning mechanism for the Sinkhorn regularization parameter ε based on local data heterogeneity. The improved system can automatically detect regions of high variability in target variables and switch to smaller ε for geometric fidelity, while using larger ε in homogeneous regions to accelerate convergence—directly leveraging the theoretical trade-off from Theorem 2.1 and 2.2.
-
Scalable Uncertainty Quantification for High-Dimensional Stochastic PDEs: Replace exact Wasserstein-based loss functions in stochastic neural networks (SNNs) with the debiased local Sinkhorn divergence. The improved system can train generative models for random fields (e.g., Darcy flow, FitzHugh–Nagumo) with significantly lower computational cost and memory footprint, enabling real-time uncertainty propagation in multidimensional systems where exact optimal transport is intractable.
-
Curse-of-Dimensionality Mitigation via Entropic Regularization: Integrate the local Sinkhorn loss as a drop-in replacement for squared W2 loss in existing probabilistic deep learning pipelines. The improved system can maintain reconstruction accuracy while achieving faster statistical convergence with fewer training samples, particularly beneficial for sparse-data scientific machine learning tasks (e.g., subsurface flow, cardiac modeling).
-
Robust Conditional Distribution Learning with Heteroscedastic Noise: Enhance SNN training to handle non-Gaussian, multimodal conditional distributions. The improved system can outperform current benchmarks (e.g., Mixture Density Networks, Conditional VAEs) by using the local Sinkhorn divergence to match full conditional distributions locally, providing more accurate mean and variance predictions in highly nonlinear stochastic systems.
-
Computationally Efficient Local Matching for Non-Stationary Random Fields: Implement a neighborhood-based decomposition that partitions the spatial/temporal domain into local patches, each optimized with the Sinkhorn divergence. The improved system can reconstruct non-stationary random fields with sharp discontinuities or localized features, achieving better geometric fidelity than global loss functions (e.g., MMD, Energy Distance) while remaining fully differentiable for end-to-end training.
-
Theoretically Grounded Hyperparameter Selection: Use the generalization bounds from Theorems 2.1 and 2.2 to guide automatic selection of ε and neighborhood size. The improved system can balance approximation bias (from neighborhood averaging) and statistical error (from empirical Sinkhorn estimation), providing principled convergence guarantees and reducing trial-and-error in model configuration.
-
Probabilistic Scientific Machine Learning for Dynamical Systems: Extend the framework to learn drift and diffusion functions in stochastic differential equations (e.g., FitzHugh–Nagumo). The improved system can train neural SDEs with local Sinkhorn divergence, yielding more accurate noise models and improved long-term predictive distributions compared to temporally decoupled W2 losses, while being computationally faster.
Abstract
In this paper, we propose a local Sinkhorn divergence framework for conditional distribution reconstruction of multidimensional random fields. By utilizing the debiased Sinkhorn divergence, our proposed approach develops a differentiable and computationally efficient local distribution matching objective to train stochastic neural networks (SNNs). Furthermore, we establish theoretical generalization error estimates for our local Sinkhorn divergence framework, which explicitly characterizes the trade-off between approximation bias and statistical efficiency controlled by the regularization parameter and reveals how our proposed local Sinkhorn divergence loss function can be efficiently applied to learning multidimensional random field models. The proposed framework provides a scalable alternative to exact local optimal transport for conditional distribution reconstruction, offering a practical compromise between geometric fidelity, statistical efficiency, and computational scalability for uncertainty quantification and probabilistic scientific machine learning. Through various numerical examples, we compare our proposed local Sinkhorn divergence framework with other loss functions to train SNNs and with other machine-learning-based uncertainty quantification frameworks, demonstrating that the proposed local Sinkhorn divergence framework achieves an effective balance between reconstruction accuracy and computational efficiency while maintaining good scalability for multidimensional stochastic systems.
Sources
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks