Conformal Data Contamination Tests for In-distribution Data Acquisition

arXiv:2507.13835 · stat.ML, cs.LG · Submitted 2025-07-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Conformal Data Contamination Tests for In-distribution Data Acquisition".

Jane: The gist The proposed work introduces a distribution-free, contamination-aware data-sharing framework that uses novel two-sample testing procedures, termed conformal data contamination tests,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So we're looking at this paper called "Conformal Data Contamination Tests for In-distribution Data Acquisition" and the authors are trying to figure out how we can share data safely without making huge assumptions about what that external data looks like.

Jane: Exactly, Tom. They're tackling the problem where you want to get more quality training data from other agents, but you have no idea if their data is actually clean or relevant for your specific task.

Lu: It's interesting because they move away from assuming every agent follows the same distribution, which is a huge simplification we often make in these types of problems.

Meng: So what's the core idea here? How do they figure out if someone is sending you bad data without knowing their internal setup?

Tom: Well, the paper proposes this new approach using something called conformal data contamination tests to identify agents whose data is most valuable for personalization. They want a distribution-free way to buy and sell training data between agents.

Jane: Think of it like this: instead of assuming everyone's cooking uses the same ingredients, they build a test that checks if someone’s ingredients are fundamentally wrong or contaminated relative to yours.

Lu: The mathematical setup they use involves decomposing the external agent distributions into a local distribution and some contamination factor, pi k, which is what we need to estimate.

Tom: Right, so the data sharing procedure they outline involves round one where agent zero gets data from everyone, computes these conformal p-values, and then decides who to keep for the next round based on those tests.

Jane: They are testing a hypothesis about whether an external agent's contamination factor is below a certain threshold pi th, and if it is, they keep that data.

Meng: That sounds like a lot of statistical machinery. How do these conformal tests actually give us confidence that we aren't just picking good data by accident?

Tom: They ground the tests in rigorous theoretical foundations for conformal outlier detection, which means the false rejection probability is controlled without making any assumptions about the underlying data distributions.

Title and authors: Lu: They introduce a class of conformal data contamination p-values for testing whether pi k is less than some threshold pi th across K different agents simultaneously.

Jane: And they give us flexibility in choosing the specific conformal score used, which lets you tailor the test to different kinds of outlier detection methods you might already be familiar with.

Tom: The results show that they derive several valid p-values for testing H zero: pi pi th using two natural test statistics, the Storey estimator and the quantile estimator <ref:2507.13835#pg1>.

Meng: Those results sound promising, but what do those specific formulas actually tell us about the performance? Are they reliable even when things get messy?

Tom: They show that asymptotically as n goes to infinity, these p-values are valid for H zero: pi pi th, using formulas like the u fisher and u sum statistics <ref:2507.13835#pg1>.

Jane: And they mention that the proposed data sharing procedure uses these tests to initialize collaborations, providing theoretical guarantees while still allowing for practical data acquisition.

Lu: The procedure involves calculating conformal scores based on received data and then selecting agents in subsequent rounds based on whether their contamination factor was rejected by the test.

Tom: So, if you had to explain this paper in one sentence right now, it's about using these distribution-free tests to strategically select data partners that are actually clean and useful for your model personalization.

Jane: It’s a framework designed to let you acquire high-quality training data while avoiding irrelevant or harmful sources by rigorously testing the contamination of external agents' data.

Meng: I see the practical application here is about building an agile trading system for data, not just theoretical math. It addresses the cost and risk involved in buying external datasets without a lot of guesswork.

Lu: The potential for this work is that it formalizes guarantees to contamination tests, which opens up possibilities for dynamic data acquisition policies in continual learning scenarios.

Tom: Let's wrap up then, Jane. We've seen how they use these conformal tests to identify valuable training data without making distributional assumptions about the external agents.

Jane: It’s a significant step because it moves quality checks from being a costly, post-hoc step after you buy data to being built right into the sharing procedure itself.

Title and authors: Tom: The paper "Conformal Data Contamination Tests for In-distribution Data Acquisition" successfully uses novel distribution-free testing procedures to detect agents whose data exceeds a contamination threshold while providing false discovery rate guarantees.

Lu: This framework allows for strategic acquisition of high-quality and personalized data while avoiding irrelevant or harmful sources.

Meng: For practical purposes, this means we can automate the decision of which external datasets are worth integrating into our training process based on statistical rigor instead of just intuition.

Jane: It’s a major win because it gives data buyers a formal way to know they're getting something trustworthy before they commit to using that external information for training.

Tom: We’ve covered how the paper uses conformal data contamination tests to rigorously select collaboration partners based on non-contamination statistics derived from two natural test statistics, u storey and u quantile.

Lu: The authors found that these Storey and Quantile conformal data contamination tests achieve the highest power in several numerical comparisons when tested against other methods.

Meng: That's a strong result; it suggests this method isn't just theoretical window dressing but has solid performance in real-world scenarios like MNIST and FEMNIST.

Jane: The paper also points out some limitations, specifically that the procedure relies on the exchangeability of the calibration and test data, which limits its direct application to time series where things change over time.

Tom: And they also noted they didn't fully consider the aspect of data subset selection in this initial work, though future work is suggested there.

Lu: The authors do propose a data-driven approach to selecting key hyperparameters like Kbudget to optimize accuracy in practice, which is a practical addition for real-world implementation.

Meng: So the big picture is that this framework provides a principled way to handle data quality uncertainty in collaborative learning environments without relying on potentially false distributional assumptions.

Jane: Ultimately, the "Conformal Data Contamination Tests for In-distribution Data Acquisition" framework gives us a robust mechanism for strategic data trading by formally guaranteeing the quality of external inputs.

Tom: That’s what they did, using conformal data contamination tests to identify agents whose data is most valuable for model personalization while providing statistical guarantees.

The paper's summary: Tom: So, we’ve been talking about this paper, "Conformal Data Contamination Tests for In-distribution Data Acquisition." Basically, they’re building a way to buy data from other sources without having to make huge assumptions about what that external data is actually like.

Jane: That's right. They introduce these conformal tests which are distribution-free, meaning you don't need to know the exact statistical recipe of the other agent's data to tell if it’s clean or contaminated.

Lu: The core mechanism they use is decomposing every external agent’s data into a local piece and some contamination factor, pi k, which is what we want to estimate.

Meng: So, instead of just guessing if someone is sending garbage, they have a rigorous test that controls the error rate without assuming anything about the data's shape.

Tom: Exactly. This lets them build an agile data trading framework where agent zero decides who to partner with next based on these contamination tests, which provides formal guarantees for the quality of that training data.

Jane: It shifts quality checking from being a messy step after you buy the data to being built directly into the acquisition process itself.

Lu: The results are quite strong because they derive valid p-values for testing whether an agent’s contamination factor is below a certain threshold, using things like Storey and Quantile estimators.

Tom: And they show that asymptotically, these specific formulas work really well as the amount of data gets bigger. They even found that the Storey and Quantile tests have the highest power in several numerical comparisons when tested against other methods.

Jane: That’s significant because it means these methods are statistically competitive with what we currently use for outlier detection, which is a big deal for practical application.

Meng: But they do mention some limitations, you know? The procedure relies on the calibration and test data being exchangeable, which means it might not work perfectly if you're dealing with time series where things constantly change.

Tom: Yeah, and they also didn't fully flesh out the data subset selection part in this initial version, though they pointed toward that as a clear next step for future work.

Jane: It’s a framework designed to let you acquire high-quality training data while avoiding irrelevant or harmful sources by rigorously testing the contamination of external agents' data.

Tom: This whole concept changes how we approach sourcing training material in collaborative AI systems, moving it from guesswork toward principled, statistically guaranteed partnerships.

Jane: So, this paper gives us a robust mechanism for strategic data trading by formally guaranteeing the quality of external inputs.

Lu: It opens up possibilities for dynamic data acquisition policies in continual learning scenarios because you can constantly test and adapt your partnership strategy based on these results.

The paper's improvements: Tom: We’ve seen how they used those conformal tests to select partners based on non-contamination statistics, but what else are they suggesting we do with this framework?

Jane: They’re talking about making it more dynamic, which means instead of just picking a fixed group for a round, you can adjust the selection strategy based on the results.

Lu: That’s where they bring in concepts like budget constraints into the selection process. They propose two main ways to choose partners: one where you stick to a fixed collaboration budget, and another that uses that data contamination information to select agents dynamically.

Meng: So it sounds like they are building a system that doesn't just pick partners once, but continuously adapts the acquisition strategy based on the quality signals they’re getting in real-time?

Tom: Exactly. If you have a set budget, you select the best candidates from within that constraint based on their contamination scores. If not, they suggest using those p-values to determine who to collaborate with next round, effectively using the contamination risk as a decision factor for future rounds.

Jane: It’s about making the data sharing smarter by letting it react to the data quality signals immediately rather than following a static plan.

Lu: This moves us closer to a truly agile trading system where your strategy evolves as you learn more about the agents' reliability, which is wild thinking for how we structure data acquisition.

Tom: That agility is key here. It formalizes how you navigate that uncertainty in exchange between agents, giving you a mathematical path forward instead of just hoping the next data batch is good.

Jane: It gives people who only listen to the show a clear picture: this isn't just about finding clean data; it’s about creating an automated, risk-aware partnership system for your AI models.

Meng: From an engineering standpoint, that means the software needs to be able to ingest those conformal scores and instantly re-calculate the optimal next round based on the latest feedback.

Lu: And Lalam sees this as a big step toward building a more trustworthy culture where data exchange is governed by rigorous quality metrics rather than just trust alone.

Tom: It’s about moving beyond simple assumptions to a principled, iterative process of data sourcing and validation.

Conclusion: Tom: So we’re wrapping up on "Conformal Data Contamination Tests for In-distribution Data Acquisition." Basically, they showed us how to use distribution-free tests to strategically select data partners based on their contamination risk.

Jane: It’s a really solid framework because it gives you statistical confidence in the quality of external training sets without needing to know the exact statistical makeup of those other agents.

Lu: The main implication is that this formalizes how we approach data sourcing, moving it from guesswork toward a principled, risk-aware partnership system for AI.

Meng: I think what this means practically is that data acquisition becomes less about blindly grabbing whatever is available and more about making calculated choices based on verifiable quality metrics.

Lalam: For culture, this advances the idea of data integrity in our systems; it builds a structure where information flow is governed by mathematical guarantees, which I find really important for building reliable AI.

Tom: That’s right. It turns the way we think about buying and selling training data between different entities.

Jane: They proved that using those conformal tests, specifically the Storey and Quantile estimators, gives you valid p-values that work well as your model size grows large.

Lu: And they found those specific tests actually have higher power than some other outlier detection methods in their comparisons on datasets like MNIST and FEMNIST.

Meng: So we’re looking at a method that is theoretically sound, practically competitive, and provides clear rules for when to trust an external data source.

Tom: It definitely offers a strong foundation for building more robust collaborative AI systems.

Jane: We’ve seen how these tests can be used to filter out agents whose data doesn't meet the required quality threshold while providing false discovery rate control.

Lu: This work opens up possibilities for dynamic acquisition policies in continual learning, allowing the system to constantly test and adapt its partner selection strategy.

Tom: That’s a big area for future research, pushing beyond just finding good partners once into continuous, adaptive data sourcing.

Department of Mathematical Sciences and Department of Electronic Systems, Aalborg University

stat.ML, cs.LG

Submitted: 2025-07-18

Updated: 2026-10-08

Importance score: 82/100

The gist: The gist The proposed work introduces a distribution-free, contamination-aware data-sharing framework that uses novel two-sample testing procedures, termed conformal data contamination tests, to

Key concepts

Data Contamination Model
This models how external agent distributions (P~k) differ from the local distribution (P0). It assumes the external data is a mixture of the local distribution and a specific outlier distribution (Pk), where πk represents how much of that outlier data is present.
Conformal Data Contamination Tests
These are rigorous statistical tests designed to check if an external agent's data exceeds a contamination threshold without making assumptions about its underlying distribution. They provide provable control over the probability of incorrectly rejecting a true null hypothesis, offering flexible outlier detection.
Data Sharing Procedure
This outlines the step-by-step process where one agent decides which other agents to acquire more data from. This decision is guided by the conformal tests to select partners who offer high-quality, non-contaminated data for improving model personalization.
p-values and FDR Control
The core output of the tests are p-values used to test hypotheses about contamination levels. These p-values are structured so that a specific procedure (Benjamini-Hochberg) can be used to control the False Discovery Rate, meaning we can manage the risk of falsely identifying contaminated data.

Terminology

Summary

The gist The proposed work introduces a distribution-free, contamination-aware data-sharing framework that uses novel two-sample testing procedures, termed conformal data contamination tests, to identify external data agents whose data is most valuable for model personalization.

Motivation and Problem

The amount of quality data in many machine learning tasks is limited to what is available locally to data owners. External data buyers need quality guarantees before purchasing, as external data may be contaminated or irrelevant to their specific learning task. Previous works primarily rely on distributional assumptions about data from different agents, relegating quality checks to post-hoc steps involving costly data valuation procedures. The focus is on developing an agile data trading framework that identifies valuable training data without any distributional assumption to buy and sell between data agents through the formalization of guarantees to data contamination tests.

Data Contamination Model and Sharing Procedure

The paper models the local distribution as P0, while external agent distributions are decomposed as P k = (1 − πk)P0 + πkPk, where Pk is a proper outlier distribution and πk ∈ [0, 1] is the contamination factor. The general data sharing procedure involves:

  1. Data agent 0 decides which data agent(s) to acquire more data from to improve personalization.

  2. The procedure is executed in rounds, where a batch of data is acquired and used to decide the policy in the next round.

  3. After termination, data agent 0 uses a data subset selection technique to filter out OOD datapoints and consequently improve personalization.

Conformal Data Contamination Tests

The core contribution is the introduction of conformal data contamination tests, which are grounded in rigorous theoretical foundations for conformal outlier detection.

** We introduce a class of conformal data contamination p-values for testing H10: πk ≤ πth, k = 1,..., K, without any distributional assumptions which provably controls the false rejection probability. The tests provide a lot of flexibility with the choice of the conformal score, and as such is compatible with a wide range of outlier detection methods. The p-value sequence for testing H10,..., HK0 is positive regression dependent on a subset (PRDS) thereby allowing for FDR control using the Benjamini-Hochberg (BH) procedure.**

Main Results and Validation

The authors derive several valid p-values for testing H0: π ≤ πth using two natural test statistics:

  1. uˆstorey, which is derived from the Storey estimator of π.

  2. uˆquantile, which is derived from the quantile estimator of π.

Asymptotically as n → ∞, the following p-values are valid for H0: π ≤ πth:

** uˆfisher = πm/th + Xm/k=1 Bmπth(k)Fχ2 2k −Tfisher − 2klog1n+1 + 2k(p1 + k/n − 1)p1 + k/n, where Tfisher is defined as-2Pm/i=1 log1n+1 − log(ˆpi).**

** uˆsum = πm/th + Xm/k=1 Bmπth(k)FIHk Tsum + k(p1 + k/n − 1)/2 p1 + k/n, where Tsum is defined as Pm/i=1 pˆi.**

The proposed data sharing procedure uses these conformal data contamination tests to initialize the collaboration of data agents through data sharing while providing theoretical guarantees. Numerical experiments on the MNIST and FEMNIST datasets validate the effectiveness of these tests and the proposed procedure, showing substantial improvements in accuracy compared to no data sharing and random baselines.

Limitations and Future Directions

The proposed data sharing procedure has limitations, including reliance on exchangeability of the calibration and test data, which limits applicability to for instance time series data. The authors also note that they did not consider the aspect of data subset selection in this work. Future work is suggested, including integration with incentive mechanisms in real-world data markets and extension to continual learning with online data acquisition policies. The paper demonstrates that the Storey and Quantile conformal data contamination tests achieve the highest power in several numerical comparisons. The authors also propose a data-driven approach to selecting key hyperparameters, such as Kbudget, to optimize accuracy in practice. The paper concludes that the proposed conformal data contamination tests are competitive with state-of-the-art tests.

Proposed Data Sharing Procedure

The procedure involves several steps for the 0-th data agent:

  1. Fit conformal score sˆ(·,(Z1,..., Zl)), and compute sˆl+1,..., sˆn.

  2. Receive data (round 1) from other data agents with Zk i m i=1, k ∈ [K].

  3. For each k ∈ [K], compute conformal scores on the test data, sˆk 1,..., sˆk m, and subsequently conformal p-values, pˆk 1,..., pˆk m using (2).

  4. Evaluate conformal non-contamination statistics Tk = T(ˆp k 1,..., pˆ k m) and find the corresponding conformal data contamination p-values, uˆk (see Section 3).

  5. If a fixed collaboration budget is given, select for collaboration in the following round data agents in Hˆ0 = σ(i: i ∈ [Kbudget]), where σ is a permutation on [K] such that Tσ(1) ≥ Tσ(2) ≥ · · · ≥ Tσ(K).

  6. Otherwise, collaborate in the following round with data agents in Hˆ0 = [K] - SBHα,γ(ˆu1,..., uˆK).

  7. Receive data (round 2,...) from other data agents with Zk m+i m i=1, k ∈ Hˆ0.

  8. Run data subset selection on all the received data.

  9. Use all local data Zi n i=1 together with the selected data to train the model, yielding f∗.

Conclusion

The framework successfully selects collaboration partners using novel distribution-free testing procedures, named conformal data contamination tests, which detects agents whose data exceed a contamination threshold while providing false discovery rate guarantees. This enables strategic acquisition of high-quality and personalized data while avoiding irrelevant or harmful data sources.

Improvements for AI systems

  1. The system can identify external data agents whose data is most valuable for model personalization by introducing a distribution-free, contamination-aware data-sharing framework. This allows for strategic acquisition of high-quality and personalized data while avoiding irrelevant or harmful sources, as outlined in Figure 1.

  2. The AI system can rigorously test null hypotheses regarding external agent contamination using novel two-sample testing procedures termed conformal data contamination tests. These tests are grounded in rigorous theoretical foundations for conformal outlier detection and are valid under arbitrary contamination levels while enabling false discovery rate control via the Benjamini-Hochberg procedure.

  3. The system can perform an agile data trading framework that identifies valuable training data without any distributional assumption, formalizing guarantees to data contamination tests to buy and sell between agents. This addresses the limitation that previous approaches are model-specific and computationally expensive and are often used post-hoc, after the data.

  4. The system can improve personalization by selecting collaboration partners based on non-contamination statistics, which is a mapping of conformal p-values to a positive real number where a large value indicates a small contamination factor, and vice versa. This leads to the selection of agents in rounds where the agent only acquire data from the data agents which was not rejected in the conformal data contamination test.

  5. The system can achieve statistically rigorous quality guarantees for aggregated external data through valid p-values derived from two natural test statistics like Tstorey and Tquantile, which are shown to be asymptotically valid as n → ∞.

Sources

Related papers