Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Security Radio. Generated commentary on the latest security and cryptography papers.
Nadia: Today's paper: "Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers".
Elias: Adversarial examples are intentionally perturbed inputs designed to alter a machine-learning model’s prediction while remaining close to the original input under a chosen perturbation constraint,
Nadia: First, who's behind it and why it matters.
Title and authors: Nadia: So, we've got this paper, "Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers," and it seems to cover both image and text domains, which is really interesting for our work here. We need to figure out exactly how these attacks work in practice and what the actual costs are for an adversary.
Elias: I agree, Nadia, because the title suggests a broad look at how noise bypasses classifiers across different types of data. It really makes you wonder what kind of inputs are easiest to manipulate and what kind of models are most vulnerable to those specific manipulations.
Priya: From a measurement standpoint, I'm curious about the scope; does this paper show us whether these evasion attacks hold true for various types of data structure or just one specific type? We need to know if these results are generalizable or just tied to the exact setup they used.
Nadia: Exactly, Priya, because we can’t just assume a result from one setup translates directly to another without seeing how the underlying mechanisms interact with different data representations. The paper seems to set up a comparison between image and text classification scenarios right from the start.
Elias: And that's where I see something important—the authors aren't treating these modalities as interchangeable; they are explicitly contrasting how robustness behaves in an image setting versus a text setting. That distinction is key for understanding the security implications.
Priya: I think that modality dependence is crucial because it tells us that we shouldn't trust a high accuracy score in one area to imply safety in another, which is something we see often when dealing with different data types.
Nadia: Right, and this paper highlights exactly that difference: the image experiment serves as a clear demonstration of evasion vulnerability, while the text experiment functions more like a controlled robustness evaluation. It sets up a really sharp contrast for us to analyze.
Elias: That contrast is what makes this paper useful because it allows us to isolate whether the vulnerability stems from the model architecture itself or from how sensitive the input representation is to specific types of noise.
Priya: When we look at those results, I’m focusing on what actually shows up in terms of performance shifts, not just the final numbers they report for clean accuracy. We need to see if those small perturbations actually cause a meaningful change in how the model interprets the data.
Title and authors: Nadia: That's fair, Priya; we aren't just looking at whether it failed or succeeded; we have to look at the magnitude of that failure and whether it's systematic or random noise. The paper lays out some pretty concrete figures showing how much accuracy drops when moving from a clean input to an adversarial one.
Elias: I’m interested in the specifics of those attacks they used, like FGSM versus PGD, because that tells us a lot about the complexity an attacker needs to employ and how much control they need over the perturbation budget.
Priya: And regarding those specific attack methods, I want to know if there are any limitations mentioned about what these attacks can or cannot achieve in terms of fooling a real-world system, which is where our work often gets complicated.
Nadia: The paper does point out that high clean accuracy isn't proof of reliability when you introduce those first-order attacks, so we have to be cautious about interpreting those initial results as a guarantee against any kind of manipulation.
Elias: That caution is exactly what I’m thinking about when we look at the mathematical assumptions underpinning these attacks; if the attack relies heavily on gradient information, then models with inherently smooth gradients might offer more resistance than we think.
Priya: I want to emphasize that even in the text experiment, where they used character substitutions and whitespace noise, they found only modest probability shifts and no actual flips from spam to ham under those specific constraints.
Nadia: That’s a really important finding because it suggests that for certain types of text manipulation, the current pipeline might actually be relatively stable when compared to the image results. It definitely points toward modality dependence being a major factor here.
Elias: Indeed, and that stability in the text setting contrasts sharply with the significant degradation we saw in the MNIST experiment under those same perturbation constraints, which tells us something fundamental about how language models handle noise differently than pixel data.
Priya: So, to summarize what I’m hearing from their results: for images, small gradients lead to big accuracy drops under PGD reaching zero point four one percent, but for text classification with those specific noise types, the changes were minor enough not to flip any spam classifications.
Title and authors: Nadia: That contrast is what really drives home the point: we have a clear evasion demonstration in one domain and a much more controlled sensitivity evaluation in the other, which is super useful for guiding how we test our own defenses.
Elias: I think the authors are pushing us toward empirical testing rather than just relying on clean performance metrics to judge security, which aligns with what we're trying to establish as a better practice in this area.
Priya: I just want to stress that the paper itself notes its limitation: the text experiment was intentionally constrained for an ethics-aware, educational scope and didn't provide automated search or deployment guidance for bypassing real-world filters.
Nadia: That’s a fair limitation to acknowledge, Priya; it keeps us grounded about what this study actually proves versus what it can't do in a practical sense. So, moving forward with this paper, we need to focus on how to build defenses that are robust across modalities rather than optimizing for just one.
Elias: It seems the paper strongly suggests that robustness behavior is highly dependent on the nature of the perturbation and whether it interacts with tokenization or semantics in a specific way. That's a deep area for cryptographic analysis as well, because we have to consider how an attacker can craft those perturbations efficiently.
Priya: I’m just hoping that this empirical approach guides future work toward developing defenses that are resilient across these distinct domains and not just tweaking the parameters for one classification task in isolation.
Nadia: That sounds like the right path forward, Priya; we need to keep pushing for testing robustness empirically rather than letting clean accuracy be our only metric for security.
Elias: So, wrapping up this discussion on "Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers," it seems the core message is that modality dictates the vulnerability profile, and robustness must be established through rigorous empirical testing against varied perturbation mechanisms.
Priya: I think that's a solid summary; it really reinforces the idea that we can't infer security just from clean accuracy scores in any domain.
Nadia: Exactly, Priya; we need to keep demanding tests that expose these vulnerabilities across different data types and perturbation techniques, not just looking for the highest clean score.
Elias: Alright team, let's take this insight into modality-dependent robustness and see what kind of defenses we can start simulating based on these findings. We’ve got some interesting ground to cover next time.
The paper's summary: Nadia: So, to kick things off, this paper boils down to showing how adversarial examples work in two completely different data types—images and text—and it really hammers home that the vulnerability profile changes drastically depending on what you're feeding the model.
Elias: That distinction is what always catches my eye; it suggests we can't just use one defense strategy for everything, which has big implications for how we design robust systems.
Priya: I’m looking at the summary and I see they set up a clear contrast between the image experiment, which shows a pretty straightforward evasion demonstration, and the text experiment, which they call a controlled robustness evaluation. That's a really important way to frame it for us as privacy researchers.
Nadia: Exactly, Priya; that framing helps us understand where we need to focus our efforts when building defenses—are we looking at pixel-level noise or semantic shifts in language?
Elias: I’m digging into the methodology they used to compare FGSM and PGD; it tells me a lot about the mathematical assumptions underlying these attacks, especially how quickly an attacker can push the perturbation budget before hitting a system boundary.
Priya: From what I'm seeing, their data really shows that for images, small gradient changes cause massive drops in accuracy, but for text under their specific noise conditions—substitutions and whitespace—the shifts are modest and don't flip spam to ham. That’s a crucial measurement they provided.
Nadia: That contrast is what makes this paper so interesting because it means we can’t just assume a model that performs well on clean data is safe; we have to test its behavior under the specific types of noise an attacker would actually use in different environments.
Elias: I think the authors are strongly arguing that robustness can't be inferred from a high clean accuracy score alone because those first-order attacks are still very effective at causing systematic degradation.
Priya: So, to wrap up their core message, the paper is pushing for us to stop treating clean performance as a guarantee of reliability and instead demand empirical testing across various perturbation mechanisms.
Nadia: That’s the central action item here; we need to design protocols that force models to prove their stability under diverse and potentially malicious inputs, rather than just checking a single metric.
Elias: This leads me to think about future work, especially how we can better quantify the difference between a simple gradient attack and something more sophisticated that might exploit specific weaknesses in the model’s architecture.
Priya: And I wonder if they'll ever expand this comparison to see how these same evasion techniques play out when models are trained on completely different data structures, like natural language versus structured data.
Nadia: That sounds like a very fertile area for our next discussion; we need to figure out how to translate these modality-specific findings into practical security requirements for deployment.
The paper's improvements: Tom: So, we’re looking at what this paper suggests we should do next to improve these evasion attacks research—basically, what the authors say needs to be done to make robustness testing better and more reliable.
Nadia: I see they are pushing for a dual-modality robustness testing protocol, which means they want us to explicitly separate our evaluations based on whether we're looking at image data or text data. That’s a smart way to stop us from overestimating how robust an AI is just because it did well in one domain.
Elias: I agree with Nadia; that separation is crucial for the cryptographer in me because it helps us see exactly which assumptions about the underlying mathematics are breaking down differently across those two modalities.
Priya: The paper suggests implementing iterative adversarial training using Projected Gradient Descent, which is a big step up from just testing single-step attacks like FGSM; it shows we need to simulate a more realistic, continuous attack budget.
Nadia: That makes sense because PGD gives us a better picture of how much effort an attacker needs to put in before the model actually fails under sustained pressure.
Elias: I’m particularly interested in the suggestion for building that controlled robustness evaluation pipeline for text, testing against a sequence of pre-defined perturbations, which is way more structured than just random noise injection.
Priya: From a measurement standpoint, I think that structured approach is excellent because it lets us precisely measure the change in spam probability and confidence trajectories under those specific constraints they outlined.
Nadia: That level of control in the text experiment is what we need to replicate if we want to really understand how models react to nuanced language manipulation rather than just random character changes.
Elias: And I see a point about testing simple preprocessing defenses, like bit-depth reduction, not just on clean data but specifically against adversarial inputs generated by stronger attacks like PGD. That’s a rigorous way to test if those defenses actually work or if they provide only partial mitigation.
Priya: That experimental setup sounds incredibly thorough for identifying where those basic defenses fall short when facing more complex evasion techniques.
Nadia: So, the overall improvement suggested is moving toward a much more empirical and controlled methodology that forces models to prove their stability across different data types and attack strategies.
Elias: This moves us away from just hoping for the best with clean accuracy and towards a verification-based approach where we actively try to break the system using known, well-defined attack vectors.
Priya: It sounds like the next phase involves designing tests that are not just about checking if a model failed, but about understanding precisely *why* and *how* it failed across different data representations.
Nadia: Exactly; we're moving from simply finding vulnerabilities to systematically characterizing them in a way that respects the difference between an image classifier and a text classifier.
Conclusion: Nadia: So, to wrap up this discussion on "Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers," we’ve seen how modality dictates vulnerability and that empirical testing is the only way to know if a model is actually safe.
Elias: That's right; the core finding is that robustness isn't a single property but something you have to measure against specific attack mechanisms across different data formats.
Priya: I’m just thinking about how this impacts privacy research, because understanding these noise bypasses helps us design better defenses against real-world manipulation in sensitive systems.
Nadia: Exactly, Priya; we're moving past simple accuracy scores and needing methods that test the system's behavior under pressure from various sources.
Elias: I think the implication for cryptography is that if an attacker can craft a perturbation that bypasses a classifier this easily, it gives us ideas on how attackers might approach verifying or manipulating encrypted data structures in the future.
Priya: It really highlights that we can't just focus on one type of data; we need to account for the fact that language models and image classifiers are fundamentally different entities when it comes to adversarial noise.
Nadia: That distinction is what makes this paper so important because it grounds our security thinking in reality, showing us exactly where the weaknesses lie in both domains.
Elias: We gotta keep asking which parameters break these attacks; if we can figure out the exact mathematical assumptions that make FGSM or PGD work better in one setting than another, we get closer to building defenses that are mathematically sound.
Priya: I just think this study provides a really solid foundation for future work where we can systematically test how different kinds of data corruption affect model integrity in a measurable way.
Nadia: It’s time to use these findings to build more rigorous testing protocols, ensuring that when we deploy AI systems, they're tested against the full spectrum of possible manipulations.
Elias: We should probably think about how this relates to those other papers on verifiable inference and prompt injection because both deal with making sure an AI is behaving according to its intended rules under stress.
Priya: I hope future research continues to focus on that cross-modality comparison, because that’s where we see the most meaningful insights into real-world reliability issues.
Minot State University
cs.CR, cs.CL, cs.LG
Submitted: 2026-09-10
Updated: 2026-09-10
Comments: Presented at the 58th Midwest Instruction and Computing Symposium (MICS 2026), Eau Claire, WI, March 27 to 28, 2026. 14 pages, 7 figures, 4 tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 72/100
The gist: Adversarial examples are intentionally perturbed inputs designed to alter a machine-learning model’s prediction while remaining close to the original input under a chosen perturbation constraint,
Key concepts
- Adversarial Examples
- These are intentionally slightly modified inputs, like adding tiny noise to a picture or text, designed specifically to trick a machine learning model into making the wrong prediction. The goal is to find the smallest possible change that causes this misclassification.
- FGSM and PGD Attacks
- These are specific mathematical methods used in image attacks. FGSM (Fast Gradient Sign Method) is a quick way to find a perturbation direction, while PGD (Projected Gradient Descent) is a more thorough, iterative method that refines the attack step-by-step to maximize the chance of success.
- Modality Dependence
- This refers to how sensitive a model's vulnerability is based on the type of data it processes. The study found that image models were highly vulnerable to these attacks, while text models were relatively stable when subjected to specific, controlled text modifications.
Terminology
Summary
Adversarial examples are intentionally perturbed inputs designed to alter a machine-learning model’s prediction while remaining close to the original input under a chosen perturbation constraint, and this study investigates how these manipulations bypass classifiers in both image and text domains. The gist is that adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation.
Experimental Design and Modality Contrast
The study employs two complementary experiments to evaluate adversarial robustness across different data modalities: image classification and text classification. The first experiment utilizes the MNIST handwritten-digit benchmark, featuring a compact convolutional neural network trained on this dataset. This setting serves as a clear and visually interpretable demonstration of adversarial vulnerability.
The second experiment shifts to text classification, fine-tuning DistilBERT on the SMS Spam Collection. This second setting is described as a controlled robustness evaluation
where the classifier's spam probability is examined under a controlled sequence of small, pre-defined text perturbations.
Image Evasion Attacks on MNIST
The image experiment focuses on evaluating two canonical white-box attacks: the Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD). The MNIST classifier achieved a clean test accuracy of 98.63%.
Under FGSM, accuracy fell significantly as the perturbation budget increased; for instance, at ε = 0.15, it dropped to 60.20%, and at the maximum tested budget of ε = 0.30, it reached only 1.72%. PGD was evaluated as a stronger iterative first-order attack baseline,
and its results were even more damaging than FGSM across most perturbation levels, showing substantial degradation at higher budgets.
Text Robustness Evaluation on DistilBERT
The text experiment examines how the classifier’s spam probability changes under a controlled sequence of pre-defined text perturbations.
The perturbations include character substitutions, whitespace noise, and a benign suffix.
Unlike the image setting where attacks caused large accuracy losses, in this text experiment, the selected surface-level changes produced only modest probability shifts in most displayed examples and no flip from spam to ham.
This suggests that the current pipeline is best interpreted as a controlled robustness or sensitivity evaluation rather than as a strong text-evasion demonstration.
Robustness Interpretation and Limitations
The findings lead to two distinct interpretations based on modality. In the MNIST benchmark, the results confirm clear successful evasion demonstration,
where small gradient-based perturbations cause large and systematic accuracy losses. Conversely, in the DistilBERT setting, it indicates that a compact transformer classifier can remain comparatively stable under mild perturbations.
The paper cautions that high clean-data performance should not be treated as evidence of robustness, especially when first-order attacks are involved. Furthermore, the study acknowledges limitations: MNIST is a simplified benchmark and the text experiment was intentionally constrained to an ethics-aware, educational scope,
lacking automated search or deployment guidance for bypassing real-world filters.
Comparative Summary
The combined results highlight that robustness behavior is modality-dependent.
The MNIST experiment demonstrates successful evasion under image conditions, while the text experiment shows stability under the specific perturbation design used. This comparison supports the conclusion that robustness must be established empirically rather than inferred from clean-data performance alone,
with its effect depending heavily on how perturbations interact with tokenization and semantics. The study concludes that high clean-data performance does not guarantee reliable behavior once inputs are exposed to malicious or carefully crafted perturbations.
The gist is that adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation.
Key Enumerated Findings:
-
MNIST Experiment: FGSM accuracy fell to 60.20% at ε = 0.15 and 1.72% at ε = 0.30; PGD fell to 32.47% and 0.41%.
-
Text Experiment: Controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed examples and no flip from spam to ham.
-
Conclusion: Robustness must be tested empirically rather than inferred from clean accuracy alone.
-
Comparative Summary: Image experiment is a clear successful evasion demonstration; text experiment is a local robustness/sensitivity study rather than a strong evasion demonstration.
References
[1] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
[2] I. J. Goodfellow, J. Shlens, and C.
Improvements for AI systems
Here are specific improvements to AI systems derived from the findings in this paper:
-
The system should incorporate a dual-modality robustness testing protocol, explicitly separating evaluation based on input modality (image vs. text). This prevents the overestimation of robustness based on performance in a single domain.
-
For image classification models, implement iterative adversarial training using Projected Gradient Descent (PGD) as the primary defense simulation, rather than relying solely on single-step attacks like FGSM for testing robustness boundaries. The system should be able to quantify the degradation difference between FGSM and PGD under identical perturbation budgets.
-
For text classification models (like DistilBERT), build a controlled robustness evaluation pipeline that tests performance against a sequence of pre-defined, surface-level perturbations (character substitutions, whitespace noise, benign suffixes). The system must be able to measure the precise change in spam probability and confidence trajectories under these specific constraints to determine if the model exhibits
controlled sensitivity
orstrong evasion.
-
Develop a rigorous methodology for evaluating simple preprocessing defenses (like bit-depth reduction) by testing their efficacy not just on clean data, but specifically on adversarial inputs generated by stronger attacks (e.g., PGD). The system should be designed to identify cases where a defense provides only partial mitigation without achieving baseline robustness levels.
-
The system should employ an empirical framework that dictates that high clean-data accuracy alone is insufficient evidence of security or reliability in any domain (image or text). It must mandate the comparison of performance across diverse perturbation mechanisms rather than relying on inferred robustness from metrics alone.
-
The system's documentation and reporting modules must be structured to include detailed, reproducible logging for all hyperparameters (learning rates, batch sizes, attack budgets like ε and PGD steps) and explicit qualitative outputs (e.g., confidence scores for specific adversarial examples), ensuring transparency required for ethical deployment in sensitive fields like content filtering or fraud detection.
Abstract
This paper presents a reproducible, educational study of evasion attacks in image classification and text classification. A compact convolutional network trained on MNIST reached 98.63% clean test accuracy and was evaluated under two white-box attacks. Under FGSM, accuracy fell to 60.20% at ε = 0.15 and 1.72% at ε = 0.30; under PGD it fell to 32.47% and 0.41%, and a bit-depth-reduction defense recovered only part of the loss. In the second experiment, DistilBERT fine-tuned on the SMS Spam Collection reached 98.75% accuracy and a 94.96% F1-score, but a controlled sequence of pre-defined perturbations (character substitutions, whitespace noise, and a benign suffix) produced only modest probability shifts in most displayed examples and no flip from spam to ham. Adversarial vulnerability is strongly modality-dependent: the MNIST experiment is a clear evasion demonstration, whereas the text experiment is a controlled robustness evaluation. Robustness must be tested empirically rather than inferred from clean accuracy.
Sources
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs