RPP: A Certified Poisoned-Sample Detection Framework for Backdoor Attacks under Dataset Imbalance
cs.CR, cs.CV, cs.LG
Submitted: 2026-01-30
Updated: 2026-01-30
Journal ref: Transactions on Machine Learning Research, 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deep neural networks are highly susceptible to backdoor attacks, yet most defense methods to date rely on balanced data, overlooking the pervasive class imbalance in real-world scenarios that can
Terminology
Abstract
Deep neural networks are highly susceptible to backdoor attacks, yet most defense methods to date rely on balanced data, overlooking the pervasive class imbalance in real-world scenarios that can amplify backdoor threats. This paper presents the first in-depth investigation of how the dataset imbalance amplifies backdoor vulnerability, showing that (i) the imbalance induces a majority-class bias that increases susceptibility and (ii) conventional defenses degrade significantly as the imbalance grows. To address this, we propose Randomized Probability Perturbation (RPP), a certified poisoned-sample detection framework that operates in a black-box setting using only model output probabilities. For any inspected sample, RPP determines whether the input has been backdoor-manipulated, while offering provable within-domain detectability guarantees and a probabilistic upper bound on the false positive rate. Extensive experiments on five benchmarks (MNIST, SVHN, CIFAR-10, TinyImageNet and ImageNet10) covering 10 backdoor attacks and 12 baseline defenses show that RPP achieves significantly higher detection accuracy than state-of-the-art defenses, particularly under dataset imbalance. RPP establishes a theoretical and practical foundation for defending against backdoor attacks in real-world environments with imbalanced data.
Sources
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency
- IBD-PSC: Input-level Backdoor Detection via Parameter-oriented Scaling Consistency
- Deep Partition Aggregation: Provable Defense against General Poisoning Attacks
- Long-tail learning via logit adjustment
- WaNet -- Imperceptible Warping-based Backdoor Attack
- Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction Consistency
- Long-Tailed Backdoor Attack Using Dynamic Data Augmentation Operations
- Label-Consistent Backdoor Attacks
- On Certifying Robustness against Backdoor Attacks via Randomized Smoothing
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs