Adversarial Robustness in Smishing Detection: A Comparative Analysis of Adversarial Fragility in Classical vs. Transformer-Based Detection Systems
Denzel Chiuseni, Athanase Bahizire, Silva Hama, Jema David Ndibwile
Carnegie Mellon University Africa
cs.CR
Submitted: 2026-08-13
Updated: 2026-08-14
Comments: 14 Pages, 1 Figure, 6 Equations, 4 Tables
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 81/100
The gist: This study evaluates the adversarial robustness of five model architectures for smishing (SMS phishing) detection: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two
Terminology
Summary
This study evaluates the adversarial robustness of five model architectures for smishing (SMS phishing) detection: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two multilingual transformers (mBERT, XLM-RoBERTa), using a dataset of 27,037 messages combining the Kaggle Swahili SMS Dataset and an English-Swahili Dataset, of which 356 (1.3%) were labelled malicious. The models are tested across three attack types—character obfuscation, structural perturbation, and cross-lingual code-switching—at three intensity levels (low, medium, high), with performance measured by the Robustness Degradation Ratio (RDR), defined as (F1 clean − F1 adversarial) / F1 clean.
Classical models are subjected to black-box generic attacks, while transformers are evaluated with attention-guided targeting, where perturbations are applied only to the top-K ranked tokens identified via attention weights from the final encoder layer (K = 3, 6, 10 for low, medium, high intensity). The clean-text baseline F1 scores are: Random Forest 0.9410, XGBoost 0.8950, CNN+BiLSTM 0.8970, mBERT 0.9429, and XLM-RoBERTa 0.9565.
The results reveal a distinct architectural boundary. Classical models experience near-catastrophic failure under character obfuscation and structural perturbation, with worst-case RDR values of 0.9533 (Random Forest, character obfuscation), 0.9877 (XGBoost, character obfuscation), 0.8943 (Random Forest, structural), 0.9145 (XGBoost, structural), and 0.7216 (CNN+BiLSTM, structural). Under code-switching, Random Forest reaches 0.7001, while XGBoost and CNN+BiLSTM degrade less (0.2486 and 0.2682, respectively). Transformers demonstrate significantly greater resilience, with worst-case RDR values of 0.1515 (mBERT, character obfuscation), 0.3221 (mBERT, structural), 0.2326 (XLM-RoBERTa, character obfuscation), and 0.3511 (XLM-RoBERTa, structural). Code-switching produces the lowest degradation for transformers, with RDR peaks of 0.0266 for mBERT and 0.0483 for XLM-RoBERTa.
Structural perturbation is the most pronounced vulnerability for transformers. Within the transformer group, XLM-RoBERTa, despite achieving a higher clean-text baseline (0.9565 vs. 0.9429), exhibits greater degradation than mBERT under high-intensity character obfuscation (0.2326 vs. 0.1515) and structural perturbation (0.3511 vs. 0.3221), demonstrating that clean-text performance is not a reliable predictor of adversarial robustness.
Statistical validation confirms these patterns are attributable to model architecture rather than sampling. A one-sided Mann-Whitney U test comparing high-intensity RDR values between classical and transformer architectures yields U = 50.0 (p = 0.0024), with classical RDR exceeding transformer RDR in 50 of 54 pairings. Cliff’s delta is 0.85, indicating a large effect size. A Friedman test per model across attack types yields χ2 = 6.0, p = 0.050 for Random Forest, XGBoost, and CNN+BiLSTM, with Kendall’s coefficient of concordance W = 1.00, confirming a consistent non-linear degradation threshold (low < medium < high) as an architectural property.
The architectural failure analysis attributes the differential fragility to structural properties of feature representation: TF-IDF models treat words as atomic tokens from exact surface forms, so homoglyph substitution collapses feature values to zero; structural perturbation shatters n-gram boundaries; and code-switching removes high-weight trigger tokens without replacement. Transformer subword tokenisation decomposes obfuscated forms into recognisable subword units, and multilingual pre-training maps cross-lingual expressions into proximate embedding regions, providing compounding resilience. The robustness inversion for XLM-RoBERTa is consistent with sharper attention concentration making decision-critical tokens more exploitable under white-box targeting.
The study concludes that clean-text performance does not predict adversarial robustness, that classical models experience near-catastrophic failure (RDR > 0.89) under character obfuscation and structural alteration, and that transformers remain substantially more resilient (RDR ≤ 0.35). The findings underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task. Proposed defences include Unicode normalisation, character n-gram augmentation, bilingual vocabulary extension, adversarial training augmentation for CNN+BiLSTM, and attention-aware adversarial training for transformers. Limitations include the small malicious-class test set (approximately 71 held-out messages), the English-Swahili-only scope, synthetic adversarial generation via the Gemini API, and the imperfect proxy of attention-based targeting.
Improvements for AI systems
Improvements to AI Systems:
-
Add adversarial robustness as a first-class evaluation metric. Train and select models not only on clean F1 but also on Robustness Degradation Ratio (RDR) across attack types, ensuring deployment candidates meet a minimum RDR threshold (e.g., ≤0.20) before release.
-
Implement Unicode normalization and homoglyph canonicalization as a preprocessing layer. Convert obfuscated characters (e.g., Cyrillic ‘а’ vs. Latin ‘a’) to a single canonical form before tokenization, reducing character-obfuscation attack surface for both classical and transformer models.
-
Augment classical TF-IDF and n-gram models with character-level features. Add character n-grams (e.g., 3- to 5-grams) alongside word-level tokens, so structural perturbations that break word boundaries still yield overlapping character patterns for classification.
-
Extend multilingual transformer vocabularies with code-switching-specific subword units. Fine-tune mBERT/XLM-RoBERTa with additional tokens for common Swahili-English code-switched phrases, improving embedding proximity for cross-lingual expressions and reducing RDR under code-switching attacks.
-
Introduce attention-aware adversarial training for transformers. During fine-tuning, apply perturbations to the top-K attention-ranked tokens (as in the paper) and add a regularization term that penalizes sharp attention concentration, making decision-critical tokens less exploitable.
-
Add adversarial training augmentation for CNN+BiLSTM. Generate synthetic adversarial examples (character obfuscation, structural perturbation, code-switching) via a generative API and mix them into the training set, improving robustness from RDR 0.72 to target ≤0.30.
-
Deploy an ensemble with architecture-specific fallback. Use a transformer (e.g., mBERT) as the primary detector for high robustness, but route inputs with high Unicode anomaly scores or low confidence to a character-level CNN that is less sensitive to structural perturbation.
-
Implement dynamic attack-intensity detection. Add a lightweight pre-classifier that estimates perturbation intensity (low/medium/high) from token-level entropy and Unicode diversity, then switches to a more conservative threshold or triggers human review when intensity is high.
What the improved AI system can do:
-
Detect smishing messages with high clean-text accuracy (F1 ≥ 0.94) while maintaining RDR ≤ 0.15 under character obfuscation, structural perturbation, and code-switching attacks, even at high intensity.
-
Correctly classify homoglyph-substituted messages (e.g., “p@yment” or Cyrillic ‘е’ in “free”) that would cause classical models to fail catastrophically.
-
Maintain stable performance on code-switched Swahili-English messages (e.g., “Tuma pesa sasa” vs. “Send pesa now”) without significant degradation, thanks to bilingual vocabulary extension.
-
Provide a confidence score that automatically flags high-intensity adversarial inputs for manual review, reducing false negatives in real-world phishing campaigns.
-
Generalize to other low-resource language pairs (e.g., Hindi-English, Arabic-French) by reusing the same adversarial training and normalization pipeline, given the architecture-agnostic nature of the proposed defences.
Abstract
Smishing detection systems are commonly trained and evaluated on clean, monolingual text. In low-resource settings, however, attackers frequently circumvent these systems through character obfuscation, cross-lingual code-switching, and structural perturbation. This study evaluates adversarial robustness for five model architectures: three classical lexical models (Random Forest, XGBoost, CNN+BiLSTM) and two multilingual transformers (mBERT, XLM-RoBERTa), using a dataset of 27,037 messages. Classical models are subjected to black-box generic attacks, while transformers are evaluated with attention-guided targeting. Each model is tested across three attack types and intensity levels, with performance measured by the Robustness Degradation Ratio (RDR). The results reveal a distinct architectural boundary: classical models experience near-catastrophic failure under character obfuscation and structural perturbation (RDR up to 0.988), whereas transformers demonstrate significantly greater resilience (RDR up to 0.351), with structural perturbation representing their most pronounced vulnerability. Effect-size analysis (Cliff's d) indicates a substantial difference between the two model categories. Within the transformer group, XLM-RoBERTa, despite achieving a higher clean-text baseline, exhibits greater degradation than mBERT. These findings demonstrate that clean-text performance is not a reliable predictor of adversarial robustness. Statistical validation using Mann-Whitney U and Friedman tests confirms that these patterns are attributable to model architecture rather than sampling. The results underscore the necessity for architecture-specific defences and frame smishing detection as an adversarial cybersecurity challenge rather than a static classification task.
Sources
- Bad Characters: Imperceptible NLP Attacks
- TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs