Evaluation of Adversarial Robustness in Arabic Language Models
Anwar Alajmi, Ayed Salman, Imtiaz Ahmad
cs.CL, cs.CR
Submitted: 2026-07-28
License: http://creativecommons.org/licenses/by/4.0/
The gist: The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities.
Terminology
Abstract
The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major security risks associated with such Natural Language Processing models is adversarial attacks. These attacks can deceive the model into the wrong prediction, raising critical model security and safety concerns. This study aims to assess the robustness of five state-of-the-art Arabic Language Models under a distinct set of Arabic adversarial attacks applied at various levels of granularity and using different example generation strategies. We also explore a defense technique based on adversarial training to enhance model robustness. The results show that insertion of diacritics can reduce the accuracy of some models by 92% while maintaining a low perturbation distance. For word-level attacks, manipulating Arabic conjunctions preserves high semantic similarity scores, low perturbation distance, and leads to an accuracy degradation of up to 58%. For sentence-level attacks, paraphrasing proves its effectiveness by an average reduction of 76% in the victim models' performance. While adversarial training improves overall resilience, with MARBERT being the most robust and AraBERT showing the greatest relative gains, challenges persist, particularly against character-level noise. These findings highlight both the potential and limitations of current defense strategies in morphologically rich languages like Arabic.
Sources
- Concrete Problems in AI Safety
- Intriguing Properties of Adversarial Examples
- Explaining and Harnessing Adversarial Examples
- Measure and Improve Robustness in NLP Models: A Survey
- Adversarial Examples in Modern Machine Learning: A Review
- Robustness May Be at Odds with Accuracy
- A Closer Look at Accuracy vs. Robustness
- AraBERT: Transformer-based Model for Arabic Language Understanding
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Reevaluating Adversarial Examples in Natural Language
- Multi-granularity Textual Adversarial Attack with Behavior Cloning
- Universal Sentence Encoder
- A Survey of Adversarial Defences and Robustness in NLP
- TextBugger: Generating Adversarial Text Against Real-world Applications
- TextAttack: A Framework for Adversarial Attacks, Data Augmentation, and Adversarial Training in NLP
- Towards Improving Adversarial Training of NLP Models
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- TextDefense: Adversarial Text Detection based on Word Importance Entropy
- Frequency-Guided Word Substitutions for Detecting Textual Adversarial Examples
- Learning to Discriminate Perturbations for Blocking Adversarial Attacks in Text Classification
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering