Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries

arXiv:2208.08697 · cs.LG, cs.CR, cs.CV · Submitted 2022-08-18 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries".

Jane: The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks,…

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So let's talk about who wrote this paper and what the title says. The paper is "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries," written by Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay, Arijit Mondal, and Partha Pratim Chakrabarti.

Jane: It’s interesting to see a team of researchers putting this specific focus on diverse decision boundaries. That sounds like they're going beyond just training models that are accurate; they’re focusing on how those models make their final decisions.

Lu: They’ve framed it as a way to create defender models that have different decision boundaries compared to the original model, which is a really neat way to think about separating the defense mechanism from the primary system.

Meng: So, if we translate that into something practical for an engineer, it means we aren't just stacking more identical defenses on top of each other; we’re trying to force them to learn different things.

Lalam: It suggests a future where security isn't just about making one big model bigger or better, but about introducing intentional differences between the components themselves.

The paper's summary: Tom: Let’s look at what they actually propose in this "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries." They suggest building an ensemble of classifiers where those classifiers have different decision boundaries relative to the original model.

Jane: So, the summary boils down to them developing two specific methods for creating these varied boundaries: one is a transformation called Split-and-Shuffle, and another is a feature restriction method called Contrast-Significant-Features.

Lu: Those two techniques are what generate the diverse gradients they’re after; they aim to ensure that when you attack the original model, the resulting adversarial examples don't transfer easily to these new defender models targeting the same class.

Meng: That’s a specific mechanism: Split-and-Shuffle splits an image into segments and shuffles them randomly to break spatial correlation in lower-level features. That sounds like it messes with the input structure itself.

Lalam: And Contrast-Significant-Features is about training a model such that the important features for the first model aren't important at all to the second model, which is a clever way to enforce feature diversity without necessarily hurting overall accuracy too much.

The paper's improvements: Tom: So what are the actual improvements they claim this methodology offers? They focus on three main contributions: developing a detection methodology that uses classifiers with diverse decision boundaries, proposing those two specific boundary design methods we just talked about, and then rigorously evaluating it.

Jane: The evaluation part is key here. They tested this against several state-of-the-art adversarial attacks, including ones that target both the original model and the detector model at the same time.

Lu: That simultaneous attack testing is crucial because it directly tests their hypothesis about how diversity helps stop transfers between models, especially when the adversary knows about both.

Meng: The results they show are interesting regarding false positives and false negatives on benchmark datasets, which tells us whether this method actually adds robustness without just creating a lot of noise in the system.

Lalam: It shows that this approach doesn't just make things harder to break; it gives us a way to measure exactly how much more robust the defense is compared to existing ensemble techniques.

Conclusion: Tom: So we’ve covered how this paper tackles adversarial attacks by creating diverse decision boundaries through Split-and-Shuffle and Contrast-Significant-Features, and they showed it works in tough experimental setups. This brings us to the wrap-up of "Resisting Adversarial Attacks in Deep Neural Networks using Diverse Decision Boundaries."

Jane: The big implication is that relying on just one robust model isn't enough; we need structural differences between our defenses to stop coordinated attacks from fooling everything at once.

Lu: It moves the focus toward designing ensembles where the models are purposefully designed to be different, not just randomly selected or trained slightly differently.

Meng: Practically, it means when we deploy these systems, we should expect a stronger defense against adversaries who have more information about our system's internals.

Lalam: I think this work points toward a future where AI security is built in by enforcing structural variety across the defense layers themselves.

Manaar Alam, Shubhajit Datta, Debdeep Mukhopadhyay, Arijit Mondal, Partha Pratim Chakrabarti

Indian Institute of Technology Kharagpur

cs.LG, cs.CR, cs.CV

Submitted: 2022-08-18

Updated: 2026-10-03

Importance score: 77/100

The gist: The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks,

Key concepts

Ensemble Defense
This defense uses multiple classifiers working together. Instead of relying on a single model's decision boundary, the ensemble combines predictions from several models trained in a way that their individual boundaries are intentionally made different. This diversity acts as a barrier against attacks.
Split-and-Shuffle
This input transformation method splits an image into multiple segments and then randomly shuffles those segments. The goal is to break the spatial correlation among low-level features in the original image, which helps create detector models with distinct decision boundaries compared to the original model.
Contrast-Significant-Features
This technique designs a second model by selectively training it so that the important features for one model are not significant for the other. By constraining parameters during training, this method ensures that the significant features used by Model 1 are different from those used by Model 2, increasing boundary diversity.

Terminology

Summary

The security of deep learning systems is an extremely important field of study as they are being deployed in several applications due to their ever-improving performance to solve challenging tasks, yet these systems remain vulnerable to crafted adversarial examples that can lead to misclassification despite being imperceptible to the human eye. This paper develops a new ensemble-based solution that constructs defender models with diverse decision boundaries with respect to the original model, aiming to reduce the chance of transferring adversarial examples from the original to the defender model targeting the same class.

The gist: The ensemble of classifiers constructed by (1) transformation of the input by a method called Split-and-Shuffle, and (2) restricting the significant features by a method called Contrast-Significant-Features are shown to result in diverse gradients with respect to adversarial attacks, which reduces the chance of transferring adversarial examples from the original to the defender model targeting the same class.<ref:2208.08697#pg2>

Motivation behind the Proposed Approach

The primary motivation is that if multiple neural network models with similar decision boundaries perform the same task, the transferability of adversarial examples makes it easier for an adversary to deceive all the models simultaneously, whereas having diverse decision boundaries makes it difficult for an adversary to deceive multiple models simultaneously <ref:2208.08697#pg2>

Our Contributions

The primary contributions are as follows:

  1. We develop a methodology for detecting adversarial perturbations using an ensemble of classifiers, where The classifiers are ensured to have diversity in the decision boundaries <ref:2208.08697#pg2>

  2. We propose two methods for designing such varying decision boundaries, namely (1) Transforming the inputs by a technique we call Split-and-Shuffle, and (2) Restricting the significant features by a method called Contrast-Significant-Features <ref:2208.08697#pg2>

  3. We evaluated the robustness of the proposed ensemble-based methodology on benchmark datasets and its effect on overall false positives and false negatives against several state-of-the-art adversarial attacks and those that target both the original model and the detector model simultaneously <ref:2208.08697#pg2>

How it works

The proposed methodology uses an ensemble of two classifiers: (1) Unprotected model MU, trained with the original dataset D, and (2) Detector model MD, trained with the same dataset D but prioritizing importance to lower-level features <ref:2208.08697#pg2>. The core idea is that MU and MD will have dissimilar decision boundaries but not significantly different accuracies, which ensures that a genuine example classified as class Ci in MU will also be classified as Ci in MD <ref:2208.08697#pg2>.

The diversity between the models is achieved through two distinct approaches for training MD:

  1. Transforming the inputs by a technique called Split-and-Shuffle, which splits an image into multiple segments and randomly shuffles all the segments to remove the spatial correlation among the lower-level features that existed in the original image <ref:2208.08697#pg2>. This transformation is shown to produce dissimilar decision boundaries <ref:2208.08697#pg2>.

  2. Restricting significant features by a method called Contrast-Significant-Features, which aims to design a new model such a way that the significant features of the first model should not be significant in second model, thereby establishing diversity between decision boundaries <ref:2208.08697#pg2>.

Building Detector Model MD with Input Transformation

The split-and-shuffle transformation is implemented via two types:

  1. Non-overlapping Transformation (T4 and T9): This divides the image into equal-sized non-overlapping segments, and the order of the shuffle is fixed for all images in a dataset <ref:2208.08697#pg2>. The resulting images are shown in Fig. 5 <ref:2208.08697#pg2>.

  2. Overlapping Transformation (T4 and T9): In this case, all the segments are extended along their sides to include more image features within the segments than the non-overlapping counterparts, which is shown in Fig. 7 <ref:2208.08697#pg2>. The overlapping transformation is observed to produce detector models with better accuracies than the non-overlapping transformations <ref:2208.08697#pg2>.

Building Detector Model MD with Diverse Feature Selection

The contrast-significant-features method establishes diversity by assigning importance to different features for different models, aiming for the significant features regarding one model for moving clean samples in the feature space to generate qpX Y r5 r1 r2 <ref:2208.08697#pg2>. The process involves:

  1. Selecting significant parameters by identifying pixels of input images that mostly affect the output of MU and then identifying parameters of hidden layers that describe those pixels <ref:2208.08697#pg2>.

  2. Training MD by restraining the significant parameters W during training, which is done by forcing the parameters in W to zero at each training epoch of model MD <ref:2208.08697#pg2>. This results in a detector model with different significant parameters, thereby incorporating diversity between decision boundaries <ref:2208.08697#pg2>.

Experimental Evaluation

The robustness is evaluated against two threat models:

Zero Knowledge Adversary (AZ)

The adversary AZ is unaware of the defense MD and generates examples for MU <ref:2208.08697#pg2>. The ensemble detects an adversarial example when it produces two different classes in both the models <ref:2208.08697#pg2>.

Perfect Knowledge Adversary (AP)

The adversary AP is aware of both MU and MD parameters and can generate examples considering both, leading to a higher chance that they can produce the same misclassification in both the models, but the perturbation will be significantly higher as the dissimilar perturbations from both the models need to be added to the input image <ref:2208.08697#pg2>.

The success rate of an adversary is defined as the portion of examples yielding the same incorrect misclassification from both the models in E <ref:2208.08697#pg2>. The results show that the attack success rate is maximum when considering only MU, and it is lower in the presence of detectors MT4D and MT9D, validating that the decision boundary of MU is more dissimilar to MT9D than MT4D <ref:2208.08697#pg2>.

The true positive rate (TPR) and false positive rate (FPR) are computed, showing that the drop in accuracy for MT9D, as shown in Table 3, increases the F P R of the proposed detection, but the ensemble MU + MT9D is more robust than the ensemble MU + MT4D against adversarial examples <ref:2208.08697#pg2>.

REFERENCES

[1] C. Szegedy et al., “Rethinking the inception architecture for computer vision,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp. 2818–2826.

[5] D. Silver et al., “Mastering the game of go with deep neural networks and tree search,” Nature, vol. 529, no. 7587, pp. 484–489, 2016.

[6] C. Szegedy et al., “Intriguing properties of neural networks,” in 2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014.

[7] I. J. Goodfellow et al., “Explaining and harnessing adversarial examples,” in 3rd International Conference on Learning Representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015.

[8] A. Kurakin et al., “Adversarial examples in the physical world,” in 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017.

[9] A. Madry et al., “Towards deep learning models resistant to adversarial attacks,” in 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018.

[10] S. Moosavi-Dezfooli et al., “Deepfool: A simple and accurate method to fool deep neural networks,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp.

Improvements for AI systems

  1. Stronger Adversarial Robustness via Diverse Decision Boundaries: The proposed ensemble approach constructs defender models with diverse decision boundaries with respect to the original model, which results in diverse gradients with respect to adversarial attacks, which reduces the chance of transferring adversarial examples from the original to the defender model targeting the same class. This means a single adversarial perturbation is less likely to successfully fool all models simultaneously.

  2. Enhanced Feature Prioritization for Detection: The method restricting significant features allows training a detector model such that the significant features of the first model should not be significant in second model, establishing diversity between decision boundaries without negatively impacting the original accuracy, as stated by the authors: The primary argument behind the proposed detection methodology is that MU and MD will have dissimilar decision boundaries but not significantly different accuracies.

  3. Robust Detection Against Strong Adversaries: The ensemble is evaluated against a stronger adversary targeting all the models within the ensemble simultaneously, demonstrating effectiveness even when an adversary has partial or complete knowledge of their internal working procedure, as opposed to only targeting the original model.

  4. Adaptive Input Transformation for Feature Correlation Removal: The use of Split-and-Shuffle transformation, particularly the overlapping version, aims to remove spatial correlation among features by splitting and shuffling images, which is intended to remove such correlation from the input images and ensure that "the dissimilar decision boundaries ensure that adversarial examples generated on a model trained with correlated features (MU) will have a different impact on the model trained by prioritizing the lower-level features (MD)."

  5. Fine-Grained Feature Selection for Detector Training: The Contrast-Significant-Features method identifies significant pixels and parameters based on gradient magnitudes from the original model, and then trains a detector model by restraining the significant parameters W, ensuring that the detector model has a distinct feature set compared to the original.

Related papers