Estimation of the sub-Gaussian Parameter

arXiv:2606.06384 · math.ST, stat.ME, stat.ML, stat.TH · Submitted 2026-06-04 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Estimation of the sub-Gaussian Parameter".

Jane: The gist The estimation of the sub-Gaussian parameter involves studying a natural estimator based on constrained maximization of an empirical analogue of L,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we’re looking at this paper by Liu and Xu titled "Estimation of the sub-Gaussian Parameter." It’s tackling something that isn't super common in statistical testing: estimating that specific parameter, xi two*, which basically defines how quickly a random variable has to decay in its tails <ref:2606.06384#pg1>.

Jane: Right, so instead of just assuming a distribution is sub-Gaussian and hoping for the best, this paper gives us a way to estimate that parameter directly using data. It’s trying to figure out that xi two* value by looking at how the empirical version of the log-generating function behaves <ref:2606.06384#pg1>.

Lu: It’s really interesting because they define this parameter as the supremum over lambda of L(lambda), which is this weighted cumulant generating function, and they are using a constrained maximization approach on the empirical analogue to get an estimator.

Meng: So, what does that actually mean for us in practice? Are we talking about getting better at testing things faster than just using standard methods?

Tom: Well, the paper shows that their natural estimator is consistent if the underlying distribution is sub-Gaussian and it has a maximizer for L, which they prove on Page one <ref:2606.06384#pg1>. They show convergence rates like O(n-one/two plus epsilon) for that, which is pretty solid <ref:2606.06384#pg1>.

Jane: That consistency hinges on this truncation gap function delta P decaying fast enough, so if that condition holds, the estimator gets very close to the true value as we get more data.

Lu: But then they hit a wall with their minimax risk analysis. They show that over *all* sub-Gaussian distributions, estimating xi two* is hard; the lower bound is (one), which means uniform consistency isn't achievable without stronger assumptions on the distribution's tail growth <ref:2606.06384#pg1,over *all* sub-Gaussian distributions>.

Tom: That’s a big caveat there, Jane. It means we can’t just throw an estimator at every possible sub-Gaussian distribution and expect it to work well for all of them simultaneously.

Jane: Exactly, and they show that imposing certain stronger assumptions on how fast the function L grows actually creates a continuum of classes where the risk bounds interpolate between (one/ n) and (one) <ref:2606.06384#pg1>.

Meng: From an engineering standpoint, that means we need to be very precise about what kind of sub-Gaussian behavior we expect from our real-world data streams. It’s not just a black box.

Lu: And when the distribution isn't even sub-Gaussian, the estimator actually diverges almost surely at a rate depending on the range of your observed data, which is useful because it acts as a natural diagnostic of mis-specification.

Tom: That’s a handy feature for debugging models where we don't know what class we're in. But what about the practical applications? They mention using this to construct p-values for Gene Ontology enrichment studies.

Jane: That’s where it gets really useful, because they suggest you can use this sub-Gaussian transform to get a reliable p-value p sG g:= (-(Z obs g - Z) squared / (two two(g))), which is an alternative to the peaks-over-threshold approach <ref:2606.06384#pg1>.

Meng: So, it’s a practical tool for large-scale permutation testing when we want a more robust way to calculate significance than just relying on those peak methods.

Lu: And under stronger assumptions—specifically if L has a unique maximizer and that maximizer is bounded—they can get an even faster convergence rate of O(n-one/two) <ref:2606.06384#pg1>. That’s when they achieve the minimax optimal rate for estimating xi two* <ref:2606.06384#pg1>.

Tom: So, we’re looking at a scenario where if our assumptions hold, we get near-optimal estimation speed. But what happens when things go wrong? They detail how the estimator diverges almost surely if the data isn't sub-Gaussian.

Jane: That divergence rate depends on the range of your observed data, like n two n C n, where C is some constant, leading to a divergence at a rate of p(n) for distributions like the Laplace distribution <ref:2606.06384#pg1>.

Meng: That divergence tells you immediately that your initial assumption about the data structure was wrong. It’s a real diagnostic tool for when things go sideways in our analysis pipeline.

Tom: And they also provide tools for inference, showing that under stronger conditions, you can construct asymptotically valid confidence intervals using quantities involving V(lambda* n).

Jane: That variance term V(lambda) is defined as four/lambda four M(two lambda) / (M(lambda) squared - two lambda psi'(lambda) +), which gives us a concrete way to build those intervals when the limiting distribution is normal.

Lu: The paper shows that while xi two* and the sub-Gaussian norm aren't equivalent—citing Leskelä and Zhukov (two thousand twenty-six)—the bounds they find for them are sharp, showing how close they can get to each other <ref:2606.06384#pg1>.

Tom: So, to sum up, this paper on "Estimation of the sub-Gaussian Parameter" gives us a method to estimate xi two* using an empirical analogue of L, with consistency rates depending on the truncation gap delta P <ref:2606.06384#pg1,Estimation of the sub-Gaussian Parameter>.

Jane: It highlights that while consistent under good conditions, achieving uniform convergence over all distributions is hard, and it provides clear divergence diagnostics when things aren't sub-Gaussian.

Meng: Ultimately, it suggests a more principled way to handle tail bounds in large-scale testing by tying the estimation directly to the underlying distribution structure.

Lu: It connects estimation directly to testing; if you can estimate xi two*, you can use it to construct p-values for things like Gene Ontology enrichment studies <ref:2606.06384#pg1>.

Tom: And that’s where we leave it for now, listeners, but keep an ear out because the next paper we look at is going to be about dynamic free-rider detection in federated learning. We’ll be talking about how that method detects clients switching roles during training.

The paper's summary: Tom: So, we're looking at this paper by Liu and Xu called "Estimation of the sub-Gaussian Parameter." It’s basically about building a way to estimate that parameter, xi two*, which tells you how fast a random variable's tails decay.

Jane: Right. The core idea is using an empirical version of the cumulant generating function to get an estimator for that parameter, and they show it works consistently under certain conditions.

Tom: And they prove consistency if the underlying distribution is sub-Gaussian and there's a good decay rate on this truncation gap function, delta P.

Jane: That’s the catch. It means if your data fits that sub-Gaussian mold, and you know how fast it decays, you can reliably estimate that parameter with enough data.

Lu: But the paper also shows that getting a consistent estimate across *all* possible sub-Gaussian distributions is actually impossible without stronger assumptions on how those functions behave.

Tom: So, even if your data looks sub-Gaussian, you still have to be careful about the exact nature of its tails for the estimation to be solid.

Jane: And they do give us a way to check if something is behaving correctly by looking at how much that parameter estimate diverges when the data isn't sub-Gaussian anymore.

Meng: It sounds like a good diagnostic tool, something you could use in real-time as data streams come in, to tell you if your assumptions about the underlying process are holding up or not.

Lalam: From an AI culture standpoint, this is really about building more robust systems that don't just guess the right model but actually understand the tail behavior of the data they are seeing.

Tom: It connects estimation directly to testing, because once you have that parameter estimate, you can use it to create reliable p-values for things like Gene Ontology enrichment studies.

Jane: That’s a practical application; it gives researchers a solid way to calculate significance in large-scale permutation tests without relying solely on the peaks-over-threshold method.

Lu: And they even provide formulas for constructing confidence intervals under stronger assumptions, which means you can build those with asymptotic validity when things line up just right.

Tom: So, the big picture is that this paper gives us a structured approach to estimating sub-Gaussian behavior, showing consistency rates and clearly pointing out where the theoretical limits lie.

Jane: It really shows that understanding the tail structure of your data isn't just academic; it’s fundamental for getting reliable results in complex statistical tests.

Meng: It reminds me how much we need those explicit diagnostics when we're building AI systems that need to operate reliably in unpredictable environments.

Lalam: Exactly, this is about making the underlying statistical assumptions transparent so the AI can make better decisions about what it sees.

Tom: Next up, we’re looking at how this estimation method fits into other areas, like how it complements methods used in those large-scale permutation tests we mentioned earlier.

The paper's improvements: Tom: So, we're looking at how this paper suggests improving their original estimation method for that sub-Gaussian parameter, xi two*.

Jane: The authors point out that they can push the convergence rate faster than what was achievable before, specifically by making sure there’s a unique maximizer for the function L.

Tom: That means if you can find one single spot where L reaches its maximum value, it helps stabilize the estimation process significantly.

Lu: And when that maximizer is bounded, it lets them achieve that faster rate of n negative half, which is pretty good news for how fast we can get accurate results.

Meng: From an engineering standpoint, a faster convergence rate means less data is needed to get a reliable estimate from the real-world measurements.

Jane: So, they're suggesting that if you can guarantee that unique maximizer exists and stays within certain bounds, you get a tighter bound on how much error there will be.

Tom: Right, and they also show how this ties into constructing those confidence intervals we talked about earlier using that variance term V.

Lu: It connects the theoretical estimation directly to a concrete way of building valid statistical intervals for inference.

Jane: That’s important because it moves the discussion from just saying "it's consistent" to actually showing you how to build trustworthy tools on top of that estimate.

Tom: It shows they aren't just looking at the estimation itself, but the whole pipeline from raw data to a reliable conclusion.

Meng: That makes sense when we think about deploying these kinds of statistical models; you need those concrete interval formulas for any real deployment scenario.

Lalam: For AI culture, this means that as we build more complex models, we need these kinds of rigorous checks built into the estimation process so the outputs are predictable and trustworthy.

Jane: It gives us a roadmap: first, find that maximizer; then use that to bound the error; then use those bounds to build your confidence intervals.

Tom: So, it’s a refinement of their original work aimed at making it more practical for real-world applications where you need speed and reliability.

Lu: And this refinement opens up avenues for exploring even more complex distributions where the initial setup might not have been as clean.

Jane: It sets the stage nicely because now we know how to get better estimates, which leads us right into what happens when your data actually starts breaking the sub-Gaussian rules.

Tom: So, we’ve seen how they improve their estimation speed, but now it’s time to look at the "what if" scenario where the data isn't behaving as expected.

Conclusion: Tom: So, to wrap up, this paper on "Estimation of the sub-Gaussian Parameter" gives us a solid way to estimate that tail parameter using empirical methods, showing consistency under good conditions and giving clear diagnostics when things go wrong.

Jane: It really shows that understanding the tail structure of your data isn't just academic; it’s fundamental for getting reliable results in complex statistical tests.

Lu: The future work they suggest is interesting because it looks at how to extend these ideas to handle even more complicated distributions where the initial setup might not be as clean.

Meng: I think the practical implication is that having these diagnostic tools for when your data isn't sub-Gaussian helps us build more stable AI systems that can operate in unpredictable environments.

Lalam: This kind of work, focusing on making statistical assumptions transparent, really contributes to a culture where we build more trustworthy models because everyone knows what the underlying assumptions are.

Tom: We’ve seen how they improve their estimation speed, but now it’s time to look at the "what if" scenario where your data isn't behaving as expected.

Jane: The main thing here is that while we can get good estimates under perfect conditions, we have clear ways to tell when those conditions are violated and what the resulting divergence looks like.

Lu: It opens up avenues for exploring even more complicated distributions where the initial setup might not have been as clean, which is a big theoretical step.

Meng: And from an engineering side, having those concrete divergence rates helps us design better error-handling mechanisms for our systems that rely on these statistical properties.

Lalam: For AI culture, this means that as we build more complex models, we need these kinds of rigorous checks built into the estimation process so the outputs are predictable and trustworthy for everyone using them.

Tom: So, in a nutshell, this paper on "Estimation of the sub-Gaussian Parameter" gives us a structured approach to estimating that tail parameter with better diagnostic tools than before.

Jane: It really shows that knowing how fast your data decays is key to getting reliable results in complex statistical tests and builds trustworthy tools on top of those estimates.

Lu: We still have a lot of theoretical space left to explore, especially regarding those more complicated distributions they mentioned for future work.

Meng: And having those concrete divergence rates helps us design better error-handling mechanisms for our systems that rely on these statistical properties in the real world.

Lalam: This kind of work, focusing on making statistical assumptions transparent, really contributes to a culture where we build more trustworthy models because everyone knows what the underlying assumptions are.

Department of Statistics, Rutgers University · Department of Genetics, Rutgers University

math.ST, stat.ME, stat.ML, stat.TH

Submitted: 2026-06-04

Updated: 2026-10-07

Code: https://github.com/LiuJ0/AMI-Benchmark

Importance score: 70/100

The gist: The gist The estimation of the sub-Gaussian parameter involves studying a natural estimator based on constrained maximization of an empirical analogue of L, proving its consistency and deriving

Key concepts

Sub-Gaussian Parameter ($\xi_2^*)$
This parameter defines how quickly the probability of extreme values in a sub-Gaussian random variable decays. It is mathematically defined as the supremum over $\lambda$ of a weighted cumulant generating function, $L(\lambda) = 2\lambda^2 \log E[e^{\lambda X}]$. A larger value means faster decay of the tail probability.
Estimator ($\hat{\xi}_{2n}$)
The proposed estimator is constructed by finding the supremum of a scaled version of $L(\lambda)$ over $\lambda$, specifically $\sup|\lambda| \le C n L(\lambda)$. The proof establishes that this estimator is consistent, meaning it converges to the true parameter as the sample size ($n$) increases, provided the underlying distribution is sub-Gaussian.
Truncation Gap Function ($\delta P$)
This function measures how close a given distribution's tail behavior is to being strictly sub-Gaussian. The control over its decay rate dictates the achievable minimax risk for estimating $\xi_2^*$. If $\delta P$ decays fast enough, better estimation rates are possible; otherwise, uniform consistency cannot be guaranteed.
Minimax Risk
This concept sets the lower bound on the best possible performance achievable by any estimator when trying to estimate $\xi_2^*$ across a class of distributions. The analysis shows that this risk is heavily influenced by the decay rate of $\delta P$. When $\gamma = 1/2$, no assumptions on $\delta P$ are needed, and uniform consistency is impossible.

Terminology

Summary

The gist The estimation of the sub-Gaussian parameter involves studying a natural estimator based on constrained maximization of an empirical analogue of L, proving its consistency and deriving minimax risk bounds that depend critically on the decay rate of a truncation gap function δP.

Definition and Motivation

The sub-Gaussian parameter ξ2∗ is defined as supλ∈R L(λ) where L(λ) = 2λ2 log Ee λX is a weighted cumulant generating function. This parameter characterizes sub-Gaussian random variables because they admit an exponential tail bound P(X ≥ t) ≤ exp− t2 / (2ξ2∗)2. The motivation for estimation comes from using empirical sub-Gaussian tail bounds in large-scale permutation tests where knowing a priori that the null distribution is sub-Gaussian allows for better power recovery.

The Estimator and Consistency

The estimator is defined as ˆξ2n = supλ≤Cn Ln(λ) for a slowly diverging Cn which can be taken as (log n)1/4. The proof shows that ˆξ2n is always consistent when the underlying distribution is subGaussian. Specifically, if there is C0 > 0 such that δ(C0) ≤ 0, then ˆξ2n − ξ2∗ = Op(n−1/2+ε) for any ε > 0.

Minimax Risk Analysis

The minimax risk of estimating ξ2∗ is governed by the decay rate of δP. For γ ∈ [0, 1/2], the minimax risk over subclasses of distributions satisfying δP (C) ≤ rγ(C) is lower bounded by Ω1/log1−2γ (n). When γ = 1/2, where r1/2(t) ≡ 1 so that no assumptions are made on δP and the subclass coincide with the class of all sub-Gaussian distributions, the minimax risk is lower bounded by Ω(1), i.e. uniform consistency is impossible.

Rate of Divergence Under Violation of Sub-Gaussianity

When the underlying distribution is not sub-Gaussian, the estimator diverges almost surely at a rate depending on the range of the observed data. Proposition 12 states that if ∆n ≥ 2 log nCn, then ˆξ2n ≥ ∆2n / (log n). For example, if X is distributed according to the mean-zero Laplace distribution, then it holds that ∆n is of order Θp(log n) so that ˆξ2n diverges at rate Ωp(log n).

Application in Gene Ontology Enrichment Studies

The estimator can be applied to Gene Ontology (GO) enrichment studies to construct p-values for a large-scale permutation test, serving as a reliable alternative to the peaks-over-threshold approach. The sub-Gaussian transform justifies the use of the p-value p sG g:= exp− (Zobsg − Z¯g)2 / (2ˆξ2n(g)).

Asymptotic Normality and Confidence Intervals

Under stronger assumptions, if L is uniquely maximized at λ∗ and L′′(λ∗) 0, and there is C0 such that δ(C0) < 0, then for any Cn → ∞, √n(ˆξ2n − ξ2∗) d−→ N(0, V (λ∗)). This allows for the construction of asymptotically valid confidence intervals using empirical estimates of the quantities involved.

Conclusion

The truncation gap δP controls the fundamental difficulty in estimating ξ2∗. Without control over δP no estimator can be uniformly consistent; stronger decay assumptions on δP lead to faster minimax rates, culminating at a rate of n−1/2 at which ˆξ2n is minimax optimal and adaptive. Under departures from sub-Gaussianity, the estimator diverges almost surely at a rate depending on the range of the observed data, providing a natural diagnostic of mis-specification. The GO enrichment application demonstrates that for large scale permutation testing, sub-Gaussian tail bounds can act as a complementary approach to the POT method. The paper concludes by noting limitations regarding the difficulty of checking hypotheses and open problems concerning the minimax lower bound.

--- Page 31 ---

A Appendix A.4 Details of Section 5 The algorithms were retrieved and implemented as described in Liu et al. (2026). The permutation test was implemented through a modification to the Empirical Pipeline (Levi et al.; 2021; Liu et al.; 2026), and the modified code can be accessed at https://github.com/LiuJ0/AMI-Benchmark.

--- Page 30 ---

A Appendix A.3 Proofs for Section 4 Before we present the proof of Proposition 11, we shall need a brief lemma.

--- Page 29 ---

A Appendix A.2 Uniform Donsker theorems The following definitions are due to van der Vaart and Wellner (2012). Let D be a normed vector space, and let B(D) denote the Borel σ-algebra on D (generated by the open sets under the norm topology of D). For a probability space (omega, F, P), we define the outer expectation for any function X: omega → R (even non-measurable) as E∗P X:= inf EP Y: Y is (F, B(R))-measurable and Y ≥ X. If D is separable, then by Theorem 1.12.1 of van der Vaart and Wellner (2012), weak convergence is equivalent to limn→∞ sup h∈BL1D E∗P h(Xn) − Eh(GP) → 0.

--- Page 30 ---

A Appendix A.1 Proofs for Section 2 Proof of Proposition 1. Firstly, since X is sub-Gaussian, M(λ) < ∞ for all λ ∈ R and hence L(λ) < ∞ as well. Now we observe that, by rearranging the inequality, we have that ξ2 ≥ supλ∈R L(λ), so ξ2∗ ≥ supλ∈R L(λ) too. Conversely, for any ε > 0 there must be λ ∈ R such that E[e λX] > exp − (ξ∗ − ε)2 / (2λ2ψ(λ) = L(λ), so that ξ2∗ is the least upper bound.

--- Page 31 ---

A Appendix A.1.1 Proof of Theorem 1 Our first main goal of this section is prove Theorem 1. First, we need the following concentration inequality. Proposition 13. Let k ∈ N0 Assume that M(λ) exists at λ, and that X ≤ B.

Improvements for AI systems

  1. The AI system can reliably estimate a sub-Gaussian parameter for mean-zero random variables in high-dimensional settings using a natural estimator derived from constrained maximization of the empirical analogue of the cumulant generating function. This estimator, based on supλ≤Cn Ln(λ) with Cn = (log n)α, provides consistency rates of Op(n−1/2+ε) or better under specific conditions on the truncation gap deltaP, allowing for robust p-value construction in large-scale permutation tests.

  2. The system can achieve minimax optimal estimation of the sub-Gaussian parameter when certain assumptions hold, specifically when arg maxL(λ) is also bounded, leading to a convergence rate of Op(n−1/2). Furthermore, under the condition that "δ(C0) < 0, the estimator achieves the optimal rate of n−1/2" for estimating ξ2+.

  3. The AI system can be used to diagnose sub-Gaussianity in data streams by observing the divergence rate of its parameter estimation; if X is a non-sub-Gaussian random variable, then ˆξ2n diverges almost surely at a rate depending on the range of the observed data, providing a natural diagnostic of mis-specification.

  4. For GO enrichment studies, the system can generate reliable p-values by applying the sub-Gaussian transform: p sG g:= exp(-(Z obs g − Z¯ g) squared / 2ˆξ2n(g)), which is a reliable alternative to the peaks-over-threshold approach, particularly in regimes where the peaks-over-threshold method is of uncertain validity.

  5. The system can perform asymptotic inference for ξ2+ by constructing asymptotically valid confidence intervals using the formula involving Vˆn(λ∗n), which relies on showing that under certain conditions, the limiting distribution is normal with variance V(λ) = 4/λ4 M(2λ) / (M(λ)2 - 2λψ'(λ) + λ2σ2 − 1.

Sources

Related papers