The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering

arXiv:2608.21262 · stat.ML, cs.LG · Submitted 2026-08-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "The Exceedance Design Effect".

Jane: As an AI researcher with a meticulous eye for detail,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So we've seen that this paper, "The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering," is fundamentally about fixing a problem where standard survey statistics fail when machine learning data is clustered. The authors argue that the traditional way we estimate sample size for averages doesn't apply here because thresholds are order statistics.

Jane: They claim that a conformal predictor, abstention gate, or safety filter promises a certain coverage rate on new data by setting a cutoff at a quantile of calibration scores. The issue is that in modern pipelines, those calibration examples aren't independent; they share prompts or documents.

Lu: The central thesis they drive home is that because of this clustering, there isn't just one effective sample size for the entire system; instead, there’s a unique sample size associated with every specific threshold level we set.

Meng: That means our initial assumptions about how many examples we need to guarantee coverage are too simplistic because the required count actually shifts depending on where we draw that line.

Lalam: Essentially, they show that the frequency with which clustered scores land on one side of a chosen threshold is what really matters, and this frequency changes based on the threshold itself.

Tom: They introduce a specific term called the Exceedance Design Effect to quantify this shift, defining it as one + (- one) rho I(p), where rho I(p) is that intra-cluster correlation of exceedance indicators at level p.

Jane: They also present a theorem showing that the limiting variance of coverage becomes proportional to p(one-p)/n eff, where the true effective sample size, n eff, is defined as being inversely proportional to this design effect.

Lu: This mathematical formulation links the design effect directly to the intra-cluster correlation of those indicator bits, making it a very precise tool for measuring the impact of clustering on thresholding systems.

Meng: So if we want to build something reliable, we can't just pick a blanket sample size; we have to account for this level-dependent effect explicitly in our calculations.

Lalam: It makes sense because it moves the focus from just looking at how similar the raw scores are numerically to looking at how often those scores actually trigger an action when compared against a specific cutoff.

Tom: The paper is really important because it shows that for tail-dependent families, this true design effect has a floor, and for mean drift scenarios, there's also a first-order bias term under clustering.

Jane: That distinction between how the effect manifests at different levels is what gives us the detailed understanding we need to move beyond naive estimations based on score correlation.

Lu: It’s a very rigorous way to model the uncertainty that arises from data dependency in these kinds of automated decision systems.

Conclusion: Tom: So, wrapping up this discussion on "The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering," the main point is that we need to stop relying on old methods that assume independence when dealing with clustered data in ML pipelines.

Jane: The authors are giving us a new way to think about sample size by showing it't not a single number but a variable dependent on which threshold we are considering, and they’ve provided the math to show how this dependency is exactly measured.

Lu: From my perspective, this provides a solid theoretical foundation for building more sophisticated uncertainty quantification tools that can handle the complexities of modern AI data structures better.

Meng: Practically speaking, it means our deployment strategies need to account for the fact that performance guarantees aren't uniform across all possible operating points of our models.

Lalam: The implication is that we have a better way to ensure the robustness of our AI outputs by looking at coverage not just as an average, but as a distribution.

Tom: Exactly. It’s about shifting from simply reporting a single number to reporting the dispersion of that coverage, which gives us a much clearer picture of what we can actually expect in deployment.

Jane: So, in simple terms, this research gives us the tools to calculate an effective sample size that respects the reality of how clustered data affects specific thresholding decisions.

Lu: It opens up possibilities for designing AI systems where uncertainty is managed with a level-specific understanding rather than a generalized, averaged approach.

Meng: For me, it means our engineering roadmap needs to integrate this kind of analysis into the validation stages to ensure we aren't overconfident in our results due to ignoring these dependencies.

Lalam: And for me, it means we can build AI systems whose cultural impact is more predictable because we are building them on a foundation that acknowledges the inherent structure of the data they consume.

stat.ML, cs.LG

Submitted: 2026-08-21

Updated: 2026-10-01

Comments: 22 pages, 2 figures. Lean proofs and code: https://doi.org/10.5281/zenodo.21595640

Code: https://github.com/ACNoonan/exceedance-design-effect

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: As an AI researcher with a meticulous eye for detail, I have thoroughly analyzed both provided texts concerning "The Exceedance Design Effect." The material presents a sophisticated critique of

Key concepts

Exceedance Design Effect
This effect measures how clustering reduces the effective sample size needed for reliable coverage guarantees. It is calculated using intra-cluster correlation of exceedance indicators, which directly reflects how often correlated scores fall on the same side of a chosen threshold.
Intra-cluster Correlation (ICC) of Exceedance Indicators
This measures the correlation between whether different data points cross a specific threshold. It is more accurate than using raw score correlations because it focuses specifically on the decision-making process—the 'exceedance'—rather than just general score similarity.
Level-Dependent Effective Sample Size ($n_{eff}$)
Unlike traditional methods, this concept asserts that there is no one universal effective sample size. Instead, $n_{eff}$ changes depending on the specific coverage level (threshold) being targeted. Reporting results per level is essential for accurate deployment.
Pooling Bias
When pooling data across different families or prompts, the resulting estimate of $n_{eff}$ is systematically too low. This happens because pooling captures common components that do not affect coverage when tested on new, independent data.

Terminology

Summary

As an AI researcher with a meticulous eye for detail, I have thoroughly analyzed both provided texts concerning The Exceedance Design Effect. The material presents a sophisticated critique of standard methods used to estimate coverage guarantees (like those in conformal predictors) when data exhibits clustering—a common occurrence in modern machine learning pipelines.

Here is a detailed, synthesized summary combining the core findings from the original paper's abstract/theorems and the practical recommendations provided in Section B.


This research addresses a critical flaw in applying standard survey statistics corrections to machine learning thresholding methods, such as conformal predictors, abstention gates, and safety filters. These methods rely on setting a cutoff based on quantiles from a calibration set (e.g., the 90th percentile for 90% coverage). While these methods promise that the stated coverage rate will hold on new data, this promise is fundamentally undermined by the reality of modern data pipelines where examples are often correlated (sharing prompts, documents, or reasoning traces), leading to clustered scores.

The central thesis is that a single effective sample size (n eff) cannot describe the performance of a thresholding system when data is clustered. Survey statistics corrections designed for independent observations (averages) are inadequate here because they fail to account for how frequently correlated scores land on the same side of a chosen threshold, and this frequency changes depending on where the threshold is set.

The paper introduces the Exceedance Design Effect, defined as:

Design Effect = 1 + (- 1) rho I(p)

where rho I(p) is the **intra-cluster correlation (ICC) of exceedance indicators at level p **. This effect is built from indicator bits rather than raw score numbers, making it directly relevant to the decision process.

Key Findings on Sample Size:

  1. Level-Dependent Sample Size: A clustered dataset possesses no single effective sample size. Instead, it has a distinct effective sample size for each specific threshold level (p) at which the system is set.

  2. Theorem 1 (Limiting Variance): Under certain conditions (b to infinity with m fixed), the limiting variance of coverage is shown to be proportional to p(1-p)/n eff, where the true effective sample size is explicitly defined as:

n eff = n over 1 + (m - 1) rho I(p)

This demonstrates that the design effect enters precisely as a reduction in the effective sample size, and this reduction is governed by rho I(p).

The research rigorously demonstrates that relying on naive methods based on score correlation (the quantity often used in existing literature) is flawed:

  • Naive Design Effect is Wrong: The naive design effect based on score correlation can overstate the dispersion penalty.

  • True Design Effect: The true design effect accurately tracks the exact standard deviation of realized coverage. It reveals that for tail-dependent families, this true design effect possesses a floor (lambda U), and for mean drift, a first-order bias term exists under clustering.

Section B translates these theoretical findings into actionable advice for deployment:

  1. Focus on the Indicator: The practitioner must estimate rho I(p) as the ICC of the exceedance indicator (i.e., whether a score crossed a specific threshold), rather than relying on correlations between raw scores.

  2. Report Dispersion, Not Just Mean: It is crucial to report both the mean coverage and its dispersion (e.g., the 5th percentile of the Beta law at n eff). The 5th percentile represents what a single deployed user actually experiences, not an average over many runs.

  3. Level-Specific Reporting: Since n eff is level-dependent, results must be reported per coverage level (alpha). A system optimized for alpha=0.10 will have a different effective sample size than one optimized for alpha=0.01.

  4. Beware of Pooling: When pooling data across families (e.g., using a base model or prompt distribution), the pooled estimate recovers the marginal ICC, which includes common components that do not affect coverage when seen by the test point. This pooling systematically understates n eff (e.g., understating it by 34% at an across-family correlation of 0.5). **Estimate within families and across families separately.

Improvements for AI systems

Here are the specific improvements and capabilities an AI system could gain by implementing the findings of this research:


The core improvement is transitioning from a naive, population-based effective sample size estimation to a threshold-specific, data-driven effective sample size that accounts for intra-cluster dependence.

  1. ​Threshold-Specific Effective Sample Size (TESS) Estimation:

Implement the derived formula for the design effect:

Effective Sample Size (neff) = n / [1 + (m - 1)ρI(p)]

This replaces the traditional, level-independent correction. The system should calculate a distinct neff value for every threshold setting chosen by the user, rather than relying on a single average value.

  1. ​Level-Dependent Confidence Bands:

Instead of reporting a fixed coverage guarantee (e.g., 90%), the system must report coverage bounds that explicitly incorporate the level-dependence of the design effect:

The system should utilize the law derived in Theorem 1, which shows that dispersion is governed by a factor dependent on both score correlation and threshold level, i.e., [1 + (m - 1)ρI(p)]. This allows for a more honest assessment of risk: coverage bands will be tighter at extreme thresholds where exceedances are rare (due to the tail attenuation quantified in §5).

  1. ​Adaptive Threshold Setting based on Correlation Structure:

The system can dynamically adjust its threshold selection strategy based on the measured intra-cluster correlation, ρI(p):

If a high correlation (e.g., ρI near 1) is detected at a specific quantile level, the system should automatically use a more conservative (larger) neff estimate for that level to ensure the stated coverage guarantee holds in practice. Conversely, if scores are less correlated at an extreme threshold, the system can use a smaller neff.

  1. ​Ragged Family Robustness:

The system must incorporate the size-biased mean cluster size substitution (Proposition 2) when dealing with real-world data where cluster sizes are informative and heterogeneous (ragged). This prevents over-correction that occurs when assuming exchangeability across all family sizes:

Use the size-biased mean cluster size, m̃ = Pj(mj 2)/Pj(mj), instead of the simple average size (m̄) in neff calculations. This ensures that the correction accurately reflects the actual data structure and avoids understating the required sample size when families are large and heterogeneous.

  1. ​Diagnostic for Data Quality (Informative Size Shift Detection):

The system should monitor for informative size shift by comparing family size statistics to score statistics:

If a change in the distribution of cluster sizes is detected alongside a shift in the mean score within those clusters, the system must flag that coverage may be biased (as described in §2.5), indicating that standard design effect corrections may be insufficient and requires further analysis using methods like target-weighted hierarchical conformal prediction.


This improved AI system can now:

  1. ​Provide a Level-Specific Risk Assessment: When deploying a threshold (e.g., 95th percentile), the system won't just give a number; it will provide the specific effective sample size needed for that exact quantile, reflecting how rare or common exceedances are at that level.

2.​Detect and Mitigate Hidden Dependence: It can identify when intra-cluster correlation is high enough to necessitate a significant increase in the required data budget (neff), providing a warning before deployment leads to under-covered predictions.

3.​Accurately Model Real Data Structure: By using the size-biased mean (m̃) for neff, the system will perform much better on heterogeneous datasets common in LLM pipelines (like those involving varying document lengths or reasoning trace depths), avoiding a conservative overestimation of data needs.

4.​Distinguish Between True Dependence and Noise: The system can differentiate between true shared-ancestry dependence (which requires the complex TESS correction) and simple conditioning on a common prompt/input, allowing for more targeted debugging of model behavior.

Abstract

Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles. We prove that grouping multiplies its large-sample variance by 1+(m-1)ρ I(p), where m is the group size, p is the target fraction, and ρ I(p) measures whether two members of a group fall on the same side of the cutoff. That correlation can differ from the correlation between the scores themselves, and it changes with the target. We give a direct proof, a counterexample to using score correlation, and an extension to unequal group sizes. A dataset therefore does not have one effective sample size. How much information it contains depends on the question you ask. In our document experiment, the same 1,000 rows carried about 217 independent observations' worth of information at the median. At the 95th percentile, they carried about 621. Nothing about the dataset changed. We asked it a different question. The number of rows is a property of the dataset. The effective sample size belongs to the analysis.

Sources

Related papers