Are Targeted Data Poisoning Attacks as Effective as We Think?
cs.LG, stat.ML
Submitted: 2025-09-08
Updated: 2026-09-14
Code: https://github.com/aks2203/poisoning-benchmark
License: http://creativecommons.org/licenses/by/4.0/
The gist: Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training.
Terminology
Abstract
Targeted data poisoning attacks manipulate model predictions on specific test samples by injecting malicious data into training. Yet existing evaluations report average attack success rates over randomly selected targets, obscuring true worst-case effectiveness. We argue that the right evaluation focuses on the hardest samples to poison. The same reasoning applies to defense: since targeted attacks leave no footprint at the distribution level, defenders should proactively identify the most vulnerable samples and apply targeted countermeasures. Given a test dataset, this paper identifies both the easiest and hardest to poison examples based on only clean model information. Specifically, we offer coarse evaluations using clean training dynamics, and fine-grained classification on poison class using poison distances and budgets. Our experiments show these metrics reliably stratify samples by poisoning vulnerability, enabling both rigorous worst-case evaluation and proactive vulnerability-aware defense.
Sources
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- Preventing Unauthorized Use of Proprietary Data: Poisoning for Secure Dataset Release
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- A Neural Algorithm of Artistic Style
- Threats to Federated Learning: A Survey
- Selective Prediction via Training Dynamics
- Very Deep Convolutional Networks for Large-Scale Image Recognition
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks