Bentkus-type asymptotic e-values
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Bentkus-type asymptotic e-values".
Tom: Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So, this paper introduces "Bentkus-type asymptotic e-values," which are presented by Diego Martinez-Taboada, Ben Chugg, and Aaditya Ramdas from Carnegie Mellon University. The core idea they’re pushing is that existing asymptotic e-values have a flaw we call the missing factor that makes their inferences too conservative.
Jane: That missing factor seems to be rooted in the fact that all current asymptotic e-values are built on an exponential function, which creates a gap between what's ideal for our threshold and what we actually use for testing, according to page one of the paper.
Lu: From a theoretical standpoint, this is significant because it means they are moving away from relying on that exponential structure in favor of something based on alpha-powered positive functions. This suggests a more flexible way to define the evidence against the null hypothesis (page two).
Meng: I'm curious how this translates into something practical for us engineers; does it just make the math cleaner, or does it actually give us better results when we’re running big simulations?
Lalam: I see a potential cultural impact here; if we can build models that naturally handle data-dependent significance levels, it could help shape how we trust and interpret complex AI outputs in real-time applications.
The paper's summary: Tom: To sum up the main points of "Bentkus-type asymptotic e-values," the authors show that their new statistics, defined as E alpha n(theta; lambda):= (Z n(theta) - lambda) alpha + I alpha(lambda), successfully eliminate that missing factor we talked about earlier.
Jane: So, in plain terms, they are proposing a new way to calculate evidence against the null hypothesis that doesn't have that scaling inefficiency inherent in the old exponential-based methods. They claim these new e-values deliver sharper inference because their delta-dependent factor is smaller than what was previously seen with exponential e-values.
Lu: The paper proves this sharpness by showing a relationship between the threshold function U delta, alpha(lambda) and an ideal baseline G-one(delta), specifically proving that G-one(delta) lambda U delta, alpha(lambda) G-one(delta/c alpha) (page two).
Meng: That inequality is pretty strong; it suggests a systematic improvement in how tightly we can bound our confidence intervals when we look at those thresholds. Does that mean less risk when deploying an AI system?
Lalam: If the inference becomes tighter, it means the confidence intervals around our predictions for complex data will be much narrower, which is a big deal for operational stability in any real-world deployment.
The paper's improvements: Tom: The authors highlight several specific improvements, focusing on how these Bentkus-type e-values perform compared to the existing ones, especially when we look at practical applications like post-hoc inference and multiple testing procedures.
Jane: They show empirically that for common levels of delta, like zero point one or zero point zero five, mixtures of these new e-values actually yield thresholds smaller than the best exponential e-value they previously found, which is a very practical result for us to see here.
Lu: The authors also established near-optimality by showing that for alpha one the threshold function U delta, alpha(lambda) is strictly convex with respect to lambda, which is a key mathematical property that allows them to prove the bounds they mentioned earlier.
Meng: When we look at multiple testing, how does this affect the speed or complexity of running these procedures in a large-scale setting? Are we adding too much computation for this theoretical gain?
Lalam: For large-scale testing, especially when using the e-BH procedure with these mixtures, they demonstrate that it uniformly improves over the exponential mixture while still controlling FDR. This suggests we can test more hypotheses effectively without resorting to overly conservative adjustments.
Conclusion: Tom: So, to wrap up this discussion on "Bentkus-type asymptotic e-values," the paper successfully connects concentration inequalities with asymptotic inference, providing a powerful alternative to exponential e-values that leads to strictly tighter inference in common settings.
Jane: Ultimately, this work gives us tools for more robust post-hoc analysis and better control over multiple testing procedures by moving away from the limitations of the missing factor. It’s about achieving sharper inference when significance levels are data-dependent.
Lu: The main implication is that we now have a family of e-values that respect arbitrary dependence and allow for post-hoc analysis under data-dependent significance levels, which is a major theoretical win for this area of statistics.
Meng: From an engineering standpoint, the practical impact is seeing tighter confidence intervals in our models and better power when testing hypotheses in complex scenarios. It’s about having more reliable statistical outputs without needing to guess conservative safety margins all the time.
Lalam: I think the most interesting cultural implication is that this mathematical refinement allows us to build AI systems whose uncertainty estimates are inherently more honest and less prone to being overly cautious in their decision-making processes.
Tom: That's a great way to put it, Lalam; we’ve really seen how these Bentkus-type asymptotic e-values offer a path toward more reliable statistical conclusions in complex AI scenarios.
Diego Martinez-Taboada, Ben Chugg, Aaditya Ramdas
Department of Statistics & Data Science · Machine Learning Department
math.ST, stat.ME, stat.ML, stat.TH
Submitted: 2026-06-04
Updated: 2026-09-29
Code: https://github.com/DMartinezT/bentkus_evalues
Importance score: 92/100
The gist: Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent.
Key concepts
- Asymptotic e-values
- These are statistical measures used for post-hoc inference and multiple testing. The paper introduces 'Bentkus-type' versions to address a missing factor in older exponential methods that made inferences too conservative.
- "Missing factor"
- 'The missing factor' refers to an issue in existing asymptotic e-values rooted in their reliance on an exponential function, which creates a gap between the ideal threshold and the actual testing level used.
- Bentkus-type asymptotic e-values
- These are new statistics defined as E alpha n(theta; lambda):= (Z n(theta) - lambda) alpha + I alpha(lambda). They are proposed to deliver sharper inference by removing the scaling inefficiency inherent in old exponential methods.
Terminology
Summary
Asymptotic e-values are emerging as a powerful alternative to asymptotic p-values, particularly in post-hoc inference and multiple testing, where significance levels may be data-dependent. This paper introduces Bentkus-type asymptotic e-values, which successfully eliminate the missing factor
inherent in existing exponential-based alternatives. By leveraging near-optimal concentration inequalities developed by Bentkus in the 2000s, the authors prove that these new e-values deliver sharper inference than existing alternatives,
leading to tighter post-hoc confidence intervals and higher rejection rates in multiple testing procedures.
Introduction and Motivation
E-values offer advantages over p-values, including validity under optional stopping and ability to combine under arbitrary dependence. However, existing asymptotic e-values suffer from the missing factor,
which results in overly conservative inference.
This issue arises because all existing asymptotic e-values are based on the exponential function, which leads to a discrepancy between the ideal threshold and the actual threshold. The paper addresses this by designing new asymptotic e-values that do not suffer from this drawback, aiming for a smaller δ-dependent factor
in their thresholds.
The Missing Factor and Existing Limitations
The missing factor is explicitly discussed using the asymptotic e-value defined by Ignatiadis et al. (2024), which involves the term E∞n(θ; λ):= exp λZn(θ) − λ2/2.
Inference based on this value depends on thresholding at level 1/δ, where the optimal threshold is related to the term p2 log(1/δ).
This contrasts with CLT-based methods like Wald confidence intervals, which depend on the δ-th quantile of the standard normal distribution. The paper notes that existing exponential e-values result in a discrepancy where "zδ < ζδ," highlighting this suboptimality.
Bentkus-type Asymptotic E-values Construction
The authors introduce a novel family of statistics called Bentkus-type asymptotic e-values, which are built upon α-powered positive functions
rather than exponentials. These e-values are defined as:
-
For a fixed λ and α ∈ [0,∞), the statistic is defined as:
-
Eαn(θ; λ):= (Zn(θ) − λ)α+ + Iα(λ), where Zn(θ) is the standardized test statistic and Iα(λ) represents the truncated moment.
The paper establishes that these statistics are near-optimal in the sense that their δ-dependent factor is at most G−1(δ/cα),
where cα is a known constant. For fixed values of δ in realistic ranges (≤ 0.1), they prove that these Bentkus-type e-values lead to tighter inference than exponential-based asymptotic e-values.
Near-Optimality and Threshold Functions
The near-optimality is formalized through the relationship between the threshold function Uδ,α(λ) and the ideal baseline G−1(δ). The paper proves that for any α ∈ [0, ∞) and δ ∈ (0, 1), it holds that:
(4.5)
G−1(δ) ≤ infλ Uδ,α(λ) ≤ G−1(δ/cα).
The authors demonstrate this by showing that the upper bound is achieved at the quantile η∗ = G−1(δ/cα), and the lower bound is established via Markov’s inequality. They also prove that for α ≥ 1, the threshold function Uδ,α(λ) is strictly convex with respect to λ.
Practical Gains in Applications
The paper demonstrates the practical superiority of Bentkus-type e-values in two key areas:
-
Post-hoc inference: The authors show that for commonly used levels of δ (e.g., 0.1, 0.05, 0.01), mixtures of Bentkus-type e-values
yield smaller thresholds than the best exponential e-value.
-
Multiple testing: When using the e-BH procedure, mixtures over optimal hyperparameters demonstrate that the Bentkus-type mixtures
uniformly improve over the exponential mixture,
leading to greater power while maintaining FDR control. The analysis shows that for sparse settings, α = 0 can outperform other families by achieving zero rejections when non-null proportions are low.
Conclusion
The work successfully connects near-optimal finite-sample concentration inequalities with asymptotic inference based on e-values. The introduction of Bentkus-type asymptotic e-values provides an alternative to exponential e-values that leads to strictly tighter inference in common settings,
resulting in greater power when combined with the e-BH procedure and tighter post-hoc confidence intervals.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to AI systems:
-
Enhance Robustness in Post-Hoc Inference for Data-Dependent Significance Levels:
-
Improve Power and FDR Control in Large-Scale Multiple Testing:
-
Develop More Reliable Confidence Sets Under Arbitrary Dependence Structures:
-
Enable More Rigorous Statistical Contract Theory Applications:
Here is a detailed breakdown of what these improved AI systems can do, drawing directly from the paper's contributions:
- Enhance Robustness in Post-Hoc Inference for Data-Dependent Significance Levels:
The system can construct confidence sets that allow the significance level to change based on the observed data (data-dependent significance levels).
Unlike traditional fixed-level inference, [the improved system] can control the worst error over all levels simultaneously,
leading to tighter post-hoc confidence intervals. This is crucial for AI applications where model uncertainty or parameter estimation requires dynamic thresholds rather than fixed margins.
- Improve Power and FDR Control in Large-Scale Multiple Testing:
The system can utilize the e-BH procedure enhanced by Bentkus-type asymptotic e-values to control the False Discovery Rate (FDR) across arbitrary dependence structures without requiring conservative corrections.
The [improved system] can increase the number of rejections while maintaining FDR control,
which is vital for high-throughput AI tasks like genomic analysis or large language model fine-tuning where thousands of hypotheses are tested simultaneously.
- Develop More Reliable Confidence Sets Under Arbitrary Dependence Structures:
The system can combine multiple e-value families (e.g., mixtures of Bentkus-type e-values) to outperform standard exponential e-values in post-hoc scenarios, especially when dealing with a small, scientifically meaningful subset of potential significance levels (e.g., testing against fixed levels like 0.1, 0.05). This makes the inference more practically relevant
by optimizing for known data-dependent targets rather than relying on overly conservative worst-case bounds.
- Enable More Rigorous Statistical Contract Theory Applications:
The system can leverage the properties of asymptotic e-values (convex combinations under arbitrary dependence) to apply them in statistical contract theory, providing stronger guarantees for inference when models or distributions are complex and dependent. This is valuable for AI systems that need to make decisions based on complex statistical constraints derived from multiple interacting data sources.
Sources
- Principal-Agent Hypothesis Testing
- Post-Hoc Large-Sample Statistical Inference
- Rao-Blackwellized e-variables
- Sharp large deviation results for sums of independent random variables
- Admissible online closed testing must employ e-values
- E-Values Expand the Scope of Conformal Prediction
- Optimal E-Values for Exponential Families: the Simple Case
- E-Values for Exponential Families: the General Case
- Asymptotic and compound e-values: multiple testing and empirical Bayes
- Post-hoc $\alpha$ Hypothesis Testing and the Post-hoc $p$-value
- Equivalence testing with data-dependent and post-hoc equivalence margins
- On the Missing Factor in Some Concentration Inequalities for Martingales
- Sharp Empirical Bernstein Bounds for the Variance of Bounded Random Variables
- Intrinsic dimension concentration inequalities for self-adjoint operators
- Bringing Closure to False Discovery Rate Control: A General Principle for Multiple Testing
Related papers
- Conformal Prediction for Dyadic Regression Under Complex Missingness
- High-Dimensional Asymptotics of Differentially Private PCA
- KL Convergence Guarantees for Score diffusion models under minimal data assumptions
- Geometric bias in eigenspace perturbation under random heterogeneous noise
- On the Asymptotic Inadmissibility of Double Machine Learning Estimators Under Structure-Agnostic Models
- Double Machine Learning of Continuous Treatment Effects with General Instrumental Variables