Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers
cs.CV, cs.CR
Submitted: 2026-09-02
Updated: 2026-09-02
Code: https://github.com/fastai/imagenette
License: http://creativecommons.org/licenses/by/4.0/
The gist: Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted
Terminology
Abstract
Deep neural networks (DNNs) are increasingly deployed in real-world vision systems, yet their predictions can be covertly manipulated by backdoor attacks, in which malicious triggers cause targeted misclassification while preserving high clean accuracy. Existing defenses often rely on model internals, training data, or clean validation samples, making them difficult to deploy when only black-box access to a trained model is available. We propose TRIM (Trigger Removal by Identifying Manipulated Regions), a deployment-oriented black-box defense that detects and selectively removes backdoor triggers at inference time without requiring model internals, training data, or clean samples. The key insight behind TRIM is to identify image regions that are responsible for anomalous model behavior and purify only those regions while preserving benign content. TRIM innovates via three key components: (i) region-based segmentation with deep feature representations, (ii) adaptive trigger discovery through inpainting and diffusion-based reconstruction to isolate regions responsible for misclassification---without assumptions about trigger type, shape, or location, and (iii) selective region purification that cleans poisoned regions while retaining benign content. To support practical deployment, TRIM further caches feature embeddings of previously identified triggers, enabling efficient recognition and avoiding redundant detection and purification. Extensive experiments across diverse datasets and backdoor types, including blended, sparse, varying-size, and multiple triggers, show that TRIM consistently outperforms existing black-box defenses, reducing attack success rates (ASR) to as low as 1.16% while preserving clean accuracy of up to 87.87%. These results demonstrate that effective backdoor mitigation is possible at inference time even when the defender has no access to any auxiliary data.
Sources
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning
- REFINE: Inversion-Free Backdoor Defense via Model Reprogramming
- Backdoor Attacks and Countermeasures on Deep Learning: A Comprehensive Review
- BadNets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain
- SCALE-UP: An Efficient Black-box Input-level Backdoor Detection via Analyzing Scaled Prediction Consistency
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks
- The "Beatrix'' Resurrections: Robust Backdoor Detection via Gram Matrices
- Label-Consistent Backdoor Attacks
- Backdoor Attack through Frequency Domain
- SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
- Faster Segment Anything: Towards Lightweight SAM for Mobile Applications
- VDC: Versatile Data Cleanser based on Visual-Linguistic Inconsistency by Multimodal Large Language Models
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models