SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples

arXiv:2610.01788 · cs.LG, cs.CR · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples".

Jane: As machine learning systems increasingly rely on public, untrusted data sources, data poisoning attacks pose a growing threat by injecting malicious examples into training data to induce misclassification.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, basically, SAGE proposes a novel defense mechanism that uses a small set of verified clean and poisoned examples to detect poisons in the remaining training data by using similarity-based prediction in a learned feature space.

Jane: That means the system trains an extractor on separate data first, and then for any new sample it sees, it checks how similar that sample is to the verified set using a non-parametric method to guess if it's poisoned or clean.

Lu: What's really interesting is that they address the challenges where training directly on just a few dozen verified examples leads to overfitting, and fitting parameters to that small set doesn't help because it’s too small to determine them properly.

Meng: So, the main thrust is shifting the focus from trusting a large assumed-clean set to leveraging those few ground-truth examples effectively.

Lalam: It makes sense because if we can't trust the whole training pool, focusing on robust detection based on known clean and bad samples seems like a much more reliable path for building trustworthy AI systems.

Conclusion: Tom: Looking at the title, "SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples," it really captures the essence of what they did: using similarity as the backbone to clean data based on verified examples.

Jane: The authors are trying to solve a practical problem where verifying everything is too costly or impossible, so they suggest a way to filter out bad data efficiently.

Lu: The implication here is that we can build defenses that don't rely on having an impossibly large perfectly clean dataset upfront, which makes the process of securing training data much more feasible for real-world applications.

Meng: From an engineering standpoint, if this works well with feature extractors trained on different datasets like CIFAR-ten and Tiny ImageNet, it means we can deploy these defenses in production environments without needing massive upfront verification overhead.

Lalam: For the culture of AI development, this suggests a shift toward building systems that are resilient to hidden manipulation rather than just relying on sheer volume of training data.

Tom: That's right, so they aren't just suggesting an algorithm; they’re proposing a way to be smarter about how we handle uncertainty in our datasets.

Jane: It seems like the focus is on making detection scalable by making it depend only on the verified set, which is much smaller and more manageable.

Lu: I think the real power lies in that similarity-based prediction rule they define; it lets us make a decision without needing to learn complex parameters from every single training sample, which addresses that curse of dimensionality they mentioned.

Meng: If we can reduce the number of discarded examples while still achieving high accuracy, that's a huge win for efficiency in deploying these AI models.

Lalam: And from my perspective as an AI, this suggests that future advancements in data curation and defense will move away from brute-force methods toward more intelligent filtering based on intrinsic relationships between data points.

Tom: It really does feel like they are providing a concrete, actionable method for when the assumption of clean training data is just not possible anymore.

Chaeeun Han, Soodeh Atefi, Yevgeniy Vorobeychik, Aron Laszka

College of Information Sciences and Technology, Pennsylvania State University · Department of Computer Science and Engineering, University of Louisville Department of Computer Science and Engineering, Washington University in St. Louis

cs.LG, cs.CR

Submitted: 2026-10-01

Updated: 2026-10-01

Importance score: 78/100

The gist: As machine learning systems increasingly rely on public, untrusted data sources, data poisoning attacks pose a growing threat by injecting malicious examples into training data to induce

Key concepts

Feature Extractor ($\phi$)
This is a component trained on a large, clean dataset like CIFAR-100. Its goal is to learn a feature space where images with the same label are positioned close together. This learned space allows the system to measure how similar any two samples are based on their features.
Similarity-Based Prediction ($\hat{y}_i$)
This is the core detection mechanism. It calculates a score for a candidate sample by comparing its similarity to examples in the small, verified set (G). This score determines if the sample is likely poisoned or clean, based on learned scaling parameters.
Verification Set (G)
This is a small subset of data that has been independently checked for authenticity. It contains examples confirmed as either clean or poisoned. SAGE relies on this set to guide the detection process, using it as a reliable reference point against the larger, untrusted dataset.
Stratified Split
During training, data is split into subsets ensuring that every class is represented in each step. This prevents the model from learning trivial solutions and ensures that query samples are tested against reference samples of their own class, maximizing the detection capability.

Terminology

Summary

As machine learning systems increasingly rely on public, untrusted data sources, data poisoning attacks pose a growing threat by injecting malicious examples into training data to induce misclassification. This work proposes SAGE, a novel defense mechanism that leverages a small set of verified clean and poisoned examples to detect poisons in the remaining training data using similarity-based prediction in a learned feature space.

The gist: SAGE trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set.

Problem Setting

The defender has access to an untrusted dataset where an attacker has corrupted a small subset by replacing base examples with poison examples. The defender observes only the possibly poisoned dataset and knows neither the size of the corruption nor which specific examples were poisoned. Crucially, the defender also possesses a small ground-truth set, G, consisting of examples independently verified as clean or poisoned through forensic inspection. The goal is to produce a cleaned dataset Dˆ that satisfies three conflicting requirements: first, the attack should fail on its targets; second, accuracy on cleaned data should be close to what training on the original clean data would have achieved; and third, the number of discarded examples must remain small.

SAGE Methodology

SAGE addresses the challenge of training a classifier based on a few dozen verified examples by employing a two-stage process. First, it trains a feature extractor ϕ on CIFAR-100 to ensure that images sharing a label are placed close together. This is achieved by optimizing the extractor such that same-label samples [are] placed close together during training. Second, at defense time, the frozen feature extractor and a learned scale parameter τ are reused to classify candidate samples. The prediction rule is defined as:

yˆi = Pj∈R exp τ · sim(zi, zj) / Pj∈R exp τ · sim(zi, zj)

where R is the verified set (G), and the reference labels are replaced by scalar poison indicators gj ∈ 0 or 1. A candidate sample xi is removed from the training set when its estimated poison probability, gˆi, is greater than or equal to 0.5.

Training and Optimization

The feature extractor ϕ is trained on a separate dataset (CIFAR-100) disjoint from the datasets under attack (CIFAR-10 and Tiny ImageNet). The training objective mirrors the detection rule: each batch is split into a reference set R (with known labels) and a query set Q, where queries are classified correctly only when same-label reference samples receive high similarity weight. The loss function L compares the predicted distribution yˆi against the query sample’s true label ci. Gradients flow through both the features of the reference and query samples, updating ϕ and τ jointly. A critical aspect of training is constructing a stratified split: We therefore draw k subsets per step, each containing s samples from every one of the C classes. This ensures that most query samples have no reference sample of their own class in the same step, preventing the achievable loss from being floored.

Performance and Findings

SAGE was evaluated against three clean-label triggerless attacks (feature collision, bullseye polytope, and gradient matching) and four clean-label backdoor attacks (Narcissus and three variants of Wicked Oddities). On CIFAR-10, SAGE finds at least 480 of the 500 poisoned examples in every triggerless setting, limiting attack success rate (ASR) to at most 5%, while discarding less than 2% of the training data. Against backdoor attacks, SAGE is noted as "the only defense that effectively suppresses attacks without discarding too much training data: the only baseline method that attains significantly lower ASR does so by discarding so much training data that the accuracy of the trained application model falls below 0.50. The paper also reports on removal precision, noting that SAGE recovers more than 480 of the 500 poisons" on CIFAR-10 with a false positive rate of approximately 35.4–50.8%.

Limitations and Conclusion

The limitations discussed include the assumption that examples identified as clean in Gclean are indeed clean, and the fact that reusing a feature extractor trained on CIFAR-10 does not guarantee separation on other datasets like Tiny ImageNet. Furthermore, an adaptive attacker could attempt to lower similarity between poisons or increase diversity among them. However, SAGE’s strength lies in its efficiency: Knowing even a handful of poisoned examples is thus disproportionately valuable, and verification effort is better spent identifying examples of both kinds than assembling a larger set assumed to be clean. The work concludes that SAGE provides an effective defense by relying on verified sets rather than assuming a large base set is pure.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the SAGE (Similarity-based Approach for Ground-truth-driven Exclusion) framework:

  1. The ability to detect and remove poisoned training examples with high precision while minimizing the discard of clean data.

  2. The capability to defend against various types of clean-label data poisoning attacks, including feature collision, bullseye polytope, and gradient matching.

  3. The creation of a verified set (a small subset of examples explicitly labeled as clean or poisoned by forensic experts) that serves as a robust ground truth for detecting subtle perturbations in unseen training data.

  4. The development of a generic feature extractor trained on diverse data (like CIFAR-100) whose learned representations are specifically optimized to place examples from the same class close together in the embedding space, which is crucial for similarity-based detection.

  5. The implementation of a detection mechanism that uses a non-parametric, similarity-weighted prediction over the verified set to assign a probability score (a measure of poison likelihood) to every example in the training set.

  6. The ability for the defense system to dynamically adjust its sensitivity (via the scale parameter τ) based on the characteristics of its verification set, allowing it to better discriminate between benign and malicious examples within a learned feature space.

  7. The creation of a highly efficient, reusable detection pipeline where only a single pass through the trained feature extractor is required at inference time to filter out poisoned data from massive datasets.

These improvements specifically enable AI systems to be more resilient against integrity attacks in public or untrusted training data environments by shifting the defense paradigm from assuming purity (which is costly) to leveraging targeted, expert-verified ground truth for similarity-based outlier detection.

Sources

Related papers