Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries".
Jane: The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks,…
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So let's talk about who wrote this paper and what the title says. The paper is "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries," written by Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay, Arijit Mondal, and Partha Pratim Chakrabarti.
Jane: It’s interesting to see a team of researchers putting this specific focus on diverse decision boundaries. That sounds like they're going beyond just training models that are accurate; they’re focusing on how those models make their final decisions.
Lu: They’ve framed it as a way to create defender models that have different decision boundaries compared to the original model, which is a really neat way to think about separating the defense mechanism from the primary system.
Meng: So, if we translate that into something practical for an engineer, it means we aren't just stacking more identical defenses on top of each other; we’re trying to force them to learn different things.
Lalam: It suggests a future where security isn't just about making one big model bigger or better, but about introducing intentional differences between the components themselves.
The paper's summary: Tom: Let’s look at what they actually propose in this "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries." They suggest building an ensemble of classifiers where those classifiers have different decision boundaries relative to the original model.
Jane: So, the summary boils down to them developing two specific methods for creating these varied boundaries: one is a transformation called Split-and-Shuffle, and another is a feature restriction method called Contrast-Significant-Features.
Lu: Those two techniques are what generate the diverse gradients they’re after; they aim to ensure that when you attack the original model, the resulting adversarial examples don't transfer easily to these new defender models targeting the same class.
Meng: That’s a specific mechanism: Split-and-Shuffle splits an image into segments and shuffles them randomly to break spatial correlation in lower-level features. That sounds like it messes with the input structure itself.
Lalam: And Contrast-Significant-Features is about training a model such that the important features for the first model aren't important at all to the second model, which is a clever way to enforce feature diversity without necessarily hurting overall accuracy too much.
The paper's improvements: Tom: So what are the actual improvements they claim this methodology offers? They focus on three main contributions: developing a detection methodology that uses classifiers with diverse decision boundaries, proposing those two specific boundary design methods we just talked about, and then rigorously evaluating it.
Jane: The evaluation part is key here. They tested this against several state-of-the-art adversarial attacks, including ones that target both the original model and the detector model at the same time.
Lu: That simultaneous attack testing is crucial because it directly tests their hypothesis about how diversity helps stop transfers between models, especially when the adversary knows about both.
Meng: The results they show are interesting regarding false positives and false negatives on benchmark datasets, which tells us whether this method actually adds robustness without just creating a lot of noise in the system.
Lalam: It shows that this approach doesn't just make things harder to break; it gives us a way to measure exactly how much more robust the defense is compared to existing ensemble techniques.
Conclusion: Tom: So we’ve covered how this paper tackles adversarial attacks by creating diverse decision boundaries through Split-and-Shuffle and Contrast-Significant-Features, and they showed it works in tough experimental setups. This brings us to the wrap-up of "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries."
Jane: The big implication is that relying on just one robust model isn't enough; we need structural differences between our defenses to stop coordinated attacks from fooling everything at once.
Lu: It moves the focus toward designing ensembles where the models are purposefully designed to be different, not just randomly selected or trained slightly differently.
Meng: Practically, it means when we deploy these systems, we should expect a stronger defense against adversaries who have more information about our system's internals.
Lalam: I think this work points toward a future where AI security is built in by enforcing structural variety across the defense layers themselves.
Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay, Arijit Mondal, Partha Pratim Chakrabarti
Indian Institute of Technology Kharagpur
cs.LG, cs.CR, cs.CV
Submitted: 2022-08-18
Updated: 2026-10-03
Importance score: 77/100
The gist: The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks,
Key concepts
- Ensemble Defense
- This defense uses multiple classifiers working together. Instead of relying on a single model's decision boundary, the ensemble combines predictions from several models trained in a way that their individual boundaries are intentionally made different. This diversity acts as a barrier against attacks.
- Split-and-Shuffle
- This input transformation method splits an image into multiple segments and then randomly shuffles those segments. The goal is to break the spatial correlation among low-level features in the original image, which helps create detector models with distinct decision boundaries compared to the original model.
- Contrast-Significant-Features
- This technique designs a second model by selectively training it so that the important features for one model are not significant for the other. By constraining parameters during training, this method ensures that the significant features used by Model 1 are different from those used by Model 2, increasing boundary diversity.
Terminology
Summary
The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks, yet these systems remain vulnerable to crafted adversarial examples that can lead to misclassification despite being imperceptible to the human eye. This paper develops a new ensemble-based solution that constructs defender models with diverse decision boundaries with respect to the original model, aiming to reduce the chance of transferring adversarial examples from the original to the defender model targeting the same class.
The gist: The ensemble of classifiers constructed by (1) transformation of the input by a method called Split-and-Shuffle, and (2) restricting the significant features by a method called Contrast-Significant-Features are shown to result in diverse gradients with respect to adversarial attacks, which reduces the chance of transferring adversarial examples from the original to the defender model targeting the same class.<ref:2208.08697#pg2>
Motivation behind the Proposed Approach
The primary motivation is that if multiple neural network models with similar decision boundaries perform the same task, the transferability of adversarial examples makes it easier for an adversary to deceive all the models simultaneously,
whereas having diverse decision boundaries makes it difficult for an adversary to deceive multiple models simultaneously
<ref:2208.08697#pg2>
Our Contributions
The primary contributions are as follows:
-
We develop a methodology for detecting adversarial perturbations using an ensemble of classifiers, where
The classifiers are ensured to have diversity in the decision boundaries
<ref:2208.08697#pg2> -
We propose two methods for designing such varying decision boundaries, namely
(1) Transforming the inputs by a technique we call Split-and-Shuffle, and (2) Restricting the significant features by a method called Contrast-Significant-Features
<ref:2208.08697#pg2> -
We evaluated the robustness of the proposed ensemble-based methodology on
benchmark datasets and its effect on overall false positives and false negatives
againstseveral state-of-the-art adversarial attacks and those that target both the original model and the detector model simultaneously
<ref:2208.08697#pg2>
How it works
The proposed methodology uses an ensemble of two classifiers: (1) Unprotected model MU, trained with the original dataset D, and (2) Detector model MD, trained with the same dataset D but prioritizing importance to lower-level features <ref:2208.08697#pg2>. The core idea is that MU and MD will have dissimilar decision boundaries but not significantly different accuracies,
which ensures that a genuine example classified as class Ci in MU will also be classified as Ci in MD
<ref:2208.08697#pg2>.
The diversity between the models is achieved through two distinct approaches for training MD:
-
Transforming the inputs by a technique called Split-and-Shuffle, which
splits an image into multiple segments and randomly shuffles all the segments to remove the spatial correlation among the lower-level features that existed in the original image
<ref:2208.08697#pg2>. This transformation is shown to producedissimilar decision boundaries
<ref:2208.08697#pg2>. -
Restricting significant features by a method called Contrast-Significant-Features, which aims to
design a new model such a way that the significant features of the first model should not be significant in second model,
thereby establishing diversity between decision boundaries <ref:2208.08697#pg2>.
Building Detector Model MD with Input Transformation
The split-and-shuffle transformation is implemented via two types:
-
Non-overlapping Transformation (T4 and T9): This divides the image into equal-sized non-overlapping segments, and the order of the shuffle is fixed for all images in a dataset <ref:2208.08697#pg2>. The resulting images are shown in Fig. 5 <ref:2208.08697#pg2>.
-
Overlapping Transformation (T4 and T9): In this case,
all the segments are extended along their sides to include more image features within the segments than the non-overlapping counterparts,
which is shown in Fig. 7 <ref:2208.08697#pg2>. The overlapping transformation is observed to producedetector models with better accuracies than the non-overlapping transformations
<ref:2208.08697#pg2>.
Building Detector Model MD with Diverse Feature Selection
The contrast-significant-features method establishes diversity by assigning importance to different features for different models, aiming for the significant features regarding one model for moving clean samples in the feature space to generate qpX Y r5 r1 r2
<ref:2208.08697#pg2>. The process involves:
-
Selecting significant parameters by identifying pixels of input images that
mostly affect the output of MU
and then identifying parameters of hidden layers that describe those pixels <ref:2208.08697#pg2>. -
Training MD by
restraining the significant parameters W during training,
which is done byforcing the parameters in W to zero at each training epoch of model MD
<ref:2208.08697#pg2>. This results in a detector model withdifferent significant parameters, thereby incorporating diversity between decision boundaries
<ref:2208.08697#pg2>.
Experimental Evaluation
The robustness is evaluated against two threat models:
Zero Knowledge Adversary (AZ)
The adversary AZ is unaware of the defense MD and generates examples for MU <ref:2208.08697#pg2>. The ensemble detects an adversarial example when it produces two different classes in both the models
<ref:2208.08697#pg2>.
Perfect Knowledge Adversary (AP)
The adversary AP is aware of both MU and MD parameters and can generate examples considering both, leading to a higher chance that they can produce the same misclassification in both the models,
but the perturbation will be significantly higher as the dissimilar perturbations from both the models need to be added to the input image
<ref:2208.08697#pg2>.
The success rate of an adversary is defined as the portion of examples yielding the same incorrect misclassification from both the models in E
<ref:2208.08697#pg2>. The results show that the attack success rate is maximum when considering only MU,
and it is lower in the presence of detectors MT4D and MT9D, validating that the decision boundary of MU is more dissimilar to MT9D than MT4D
<ref:2208.08697#pg2>.
The true positive rate (TPR) and false positive rate (FPR) are computed, showing that the drop in accuracy for MT9D, as shown in Table 3, increases the F P R of the proposed detection,
but the ensemble MU + MT9D is more robust than the ensemble MU + MT4D against adversarial examples
<ref:2208.08697#pg2>.
REFERENCES
[1] C. Szegedy et al., “Rethinking the inception architecture for computer vision,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 2818–2826.
[5] D. Silver et al., “Mastering the game of go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, 2016.
[6] C. Szegedy et al., “Intriguing properties of neural networks,” in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014.
[7] I. J. Goodfellow et al., “Explaining and harnessing adversarial examples,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015.
[8] A. Kurakin et al., “Adversarial examples in the physical world,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017.
[9] A. Madry et al., “Towards deep learning models resistant to adversarial attacks,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018.
[10] S. Moosavi-Dezfooli et al., “Deepfool: A simple and accurate method to fool deep neural networks,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp.
Improvements for AI systems
-
Stronger Adversarial Robustness via Diverse Decision Boundaries: The proposed ensemble approach constructs
defender models with diverse decision boundaries with respect to the original model,
which results indiverse gradients with respect to adversarial attacks, which reduces the chance of transferring adversarial examples from the original to the defender model targeting the same class.
This means a single adversarial perturbation is less likely to successfully fool all models simultaneously. -
Enhanced Feature Prioritization for Detection: The method
restricting significant features
allows training a detector model such thatthe significant features of the first model should not be significant in second model,
establishing diversity between decision boundaries without negatively impacting the original accuracy, as stated by the authors:The primary argument behind the proposed detection methodology is that MU and MD will have dissimilar decision boundaries but not significantly different accuracies.
-
Robust Detection Against Strong Adversaries: The ensemble is evaluated against a
stronger adversary targeting all the models within the ensemble simultaneously,
demonstrating effectiveness even when an adversary haspartial or complete knowledge of their internal working procedure,
as opposed to only targeting the original model. -
Adaptive Input Transformation for Feature Correlation Removal: The use of
Split-and-Shuffle
transformation, particularly the overlapping version, aims to remove spatial correlation among features by splitting and shuffling images, which is intended toremove such correlation from the input images
and ensure that "the dissimilar decision boundaries ensure that adversarial examples generated on a model trained with correlated features (MU) will have a different impact on the model trained by prioritizing the lower-level features (MD)." -
Fine-Grained Feature Selection for Detector Training: The
Contrast-Significant-Features
method identifiessignificant pixels and parameters
based on gradient magnitudes from the original model, and then trains a detector model byrestraining the significant parameters W,
ensuring that the detector model has a distinct feature set compared to the original.
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks