Tensor-normal maximum likelihood estimation at the operator-norm sample threshold
Hengzhi He, Guang Cheng
University of California, Los Angeles
math.ST, stat.ML, stat.TH
Submitted: 2026-08-22
Updated: 2026-08-25
Comments: Work in Progress
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: Let X 1,, X n be independent Gaussian tensors in R d 1 R d k whose covariance is a Kronecker product of k unknown positive-definite factors, and put D = product a=1 k d a and d = a d a.
Terminology
Summary
Let X 1,, X n be independent Gaussian tensors in R d 1 R d k whose covariance is a Kronecker product of k unknown positive-definite factors, and put D = product a=1 k d a and d = a d a. A recent result of Franks et al. (2026) established condition-number-free nonasymptotic guarantees for the tensor-normal maximum likelihood estimator under the sample threshold nD k squared d cubed. They asked whether the cubic dependence on d could be replaced by the operator-norm scale d squared. We answer this question affirmatively. We prove that, for t at least 1, the maximum likelihood estimator exists uniquely with high probability whenever nD at least C k squared d squared t squared, and satisfies
[
d FR(,) at most C t sqrt k d over sqrt n,
d FR(a, a) at most C t sqrt k d a d over sqrt nD.
]
For every mode of largest dimension, we also obtain the sharp Thompson bound
[
d op(a, a) at most C t d over sqrt nD.
]
No sparsity, condition-number bound or warm start is assumed. For fixed k, the threshold has the information-theoretically optimal dependence on d, and the displayed rates for the full precision and the largest factor match Gaussian minimax lower bounds up to a factor sqrt k. The proof extends a random Gram bound for local group-orbit directions to the full local Lie algebra, transports it to a fixed Thompson ball by exact conjugation, and combines sensitivity of a constrained maximum likelihood estimator with an equivariant Kirszbraun extension and Gaussian concentration. This removes the Frobenius-to-operator loss responsible for the previous extra factor d and resolves the explicit open problem posed in the earlier work. The proof was developed with substantial assistance from GPT-5.6 and Claude Fable 5, and was manually checked by the authors.
Improvements for AI systems
Based on this paper, I can improve AI systems in the following specific ways:
-
Implement a new sample-complexity threshold for tensor-normal MLE: The AI can now guarantee unique maximum likelihood estimation with high probability when nD at least C k squared d squared t squared, improving from the previous d cubed dependence. This allows the system to operate with fewer samples (by a factor of d) in high-dimensional tensor settings, making it feasible for larger tensor dimensions.
-
Achieve sharper statistical rates for factor estimation: The AI can now provide error bounds of d FR(a, a) at most C t sqrt k d a d over sqrt nD and operator-norm bounds d op(a, a) at most C t d over sqrt nD for the largest mode. This enables the system to produce more accurate estimates of individual Kronecker factors, particularly useful in applications like neuroimaging or spatiotemporal modeling where per-mode precision matters.
-
Remove reliance on condition-number bounds and warm starts: The improved AI can now run tensor-normal MLE without requiring prior knowledge of the covariance's conditioning or an initial guess close to the true parameter. This makes the system robust to ill-conditioned data and allows it to be deployed in
cold-start
scenarios where no prior estimate exists. -
**Achieve information-theoretically optimal sample thresholds for fixed k **: For a fixed number of tensor modes, the AI now operates at the minimax-optimal sample complexity with respect to d, meaning it cannot be improved further in that dimension. This allows the system to be used as a benchmark or baseline for other tensor estimation methods.
-
Provide nonasymptotic, high-probability guarantees without sparsity assumptions: The AI can now deliver finite-sample guarantees (with probability at least 1 - e-t squared) for dense tensor data, which is critical for safety-critical applications where worst-case behavior must be controlled, not just asymptotic averages.
-
Extend the random Gram bound technique to full Lie algebra: The AI can now handle local group-orbit directions in the full Lie algebra (not just a subset), enabling it to analyze more complex tensor decompositions and rotations. This is useful for systems that need to estimate parameters under arbitrary orthogonal or invertible transformations.
-
Combine exact conjugation with Thompson-ball transport for stability: The AI can now transport error bounds from a local neighborhood to a fixed Thompson ball using exact conjugation, which improves numerical stability and allows the system to maintain tight error control even when the true parameter is far from the initial point.
-
Integrate equivariant Kirszbraun extension for constrained MLE sensitivity: The AI can now use an equivariant Lipschitz extension to bound the sensitivity of the constrained maximum likelihood estimator, which is a new mathematical tool that can be reused in other estimation problems involving group-invariant losses or constraints.
What the improved AI system can do:
-
Estimate the Kronecker factors of a tensor-normal distribution from data with O(d 2) samples instead of O(d 3), enabling analysis of tensors with modes up to thousands of dimensions on limited data.
-
Provide per-mode confidence intervals with optimal scaling, allowing practitioners to identify which tensor dimensions are most uncertain.
-
Run without any initialization or preconditioning, making it deployable in automated pipelines where manual tuning is impossible.
-
Guarantee performance in the worst case (high probability) even for dense, ill-conditioned tensors, which is essential for medical imaging, genomics, and spatiotemporal climate data.
-
Serve as a drop-in replacement for existing tensor normal estimators in software libraries, automatically improving their sample efficiency and error rates.
Sources
Related papers
- Conformal Prediction for Dyadic Regression Under Complex Missingness
- Bentkus-type asymptotic e-values
- High-Dimensional Asymptotics of Differentially Private PCA
- KL Convergence Guarantees for Score diffusion models under minimal data assumptions
- Geometric bias in eigenspace perturbation under random heterogeneous noise
- On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models