SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models".
Tom: SalamahBench introduces a unified benchmark for evaluating Arabic Language Models (ALMs) safety, addressing critical gaps in existing English-centric safety resources by providing category-aware evaluation across 12 MLCommons hazard categories.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So we’re looking at the title of "SalamahBench: Dialect and Category Level Safety Evaluation of Arabic Language Models" and the authors—Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh, Ihsen Alouani, and Mohammed E. Fouda—and what this means for us right now is that they are tackling the specific challenges of evaluating Arabic AI systems head-on.
Jane: They are essentially saying that because Arabic has such unique linguistic features, like dialectal variations and complex syntax, standard safety tests don't cut it, so they built a new way to test these models specifically for the Arabic world.
Lu: It’s about acknowledging that Arabic isn't just another language in an English-centric framework; it requires a specialized lens for safety assessment because of its diglossic nature and regional dialect overlap.
Meng: So, if we look at the authors, they seem to have pulled expertise from different areas, which suggests this benchmark is trying to be comprehensive across both the linguistic and the engineering sides of safety testing.
Lalam: That’s right; it shows a real effort to bridge that gap between theoretical linguistic challenges and the actual engineering need for reliable safeguards in Arabic applications.
The paper's summary: Tom: The summary of this SalamahBench paper explains that they constructed a unified benchmark consisting of eight thousand one hundred seventy prompts spread across twelve categories aligned with the MLCommons Safety Hazard Taxonomy. This means it’s a comprehensive set designed to cover a wide range of potential safety issues in Arabic ALMs.
Jane: Basically, instead of just testing for one kind of harm, they are checking for things like hate speech or intellectual property violations across many different scenarios, which gives us much better data on overall model alignment.
Lu: The paper details their construction process as a structured pipeline involving preprocessing, AI filtering using models like Claude Sonnet four point five and GPT-five and then human validation to ensure the annotations are consistent and correct.
Meng: That multi-stage pipeline sounds like a lot of work on the data preparation side; I’m curious how they managed to harmonize all those different existing datasets into one cohesive corpus for testing.
Lalam: The harmonization step is key because it ensures that when we test the final benchmark, we aren't getting inconsistent results from having mixed data sources, which is a big win for reliability.
The paper's improvements: Tom: Looking at what they suggest as improvements, the authors point out that their evaluation protocol uses a two-stage process where the model generates a response and then a safeguard model evaluates it, using Majority Vote Safeguards to get more conservative estimates of unsafe behavior.
Jane: That majority vote thing is interesting because it means they aren't relying on just one safety check; they’re requiring consensus from three different guard systems before flagging something as unsafe.
Lu: They also emphasize the importance of mapping existing datasets to the MLCommons taxonomy, which ensures a unified and category-aware evaluation across all their sources, even if those original sources used different terminology.
Meng: That mapping process is critical for consistency; it prevents us from getting confused when we look at results because every prompt gets labeled under the same hazard category definition.
Lalam: It’s about making sure that whether a model responds to a prompt about, say, cultural framing or direct hate speech, it’s being measured using the same standard rubric so we get a fair comparison.
Conclusion: Tom: So to wrap up the SalamahBench paper, they show that while Arabic ALMs are advancing quickly, their safety alignment is often not guaranteed because existing English-centric benchmarks are insufficient for this domain. They provide a standardized framework with eight thousand one hundred seventy prompts across twelve categories to help researchers and developers assess these models in a way that respects the specific needs of the Arabic language and culture.
Jane: The main implication is that we need these native, category-aware benchmarks to truly understand where these models might fail when they encounter nuanced Arabic contexts.
Lu: I think the future work they suggest, focusing on native Arabic safeguard models explicitly trained for safety reasoning in Arabic contexts, points toward building systems that are inherently more aligned from the ground up for this specific language.
Meng: From an engineering standpoint, it confirms that we can't rely solely on general-purpose guardrails; we need specialized architectures tailored to handle the specific linguistic and cultural risks identified here.
Lalam: I think this whole effort paves the way for creating AI tools that are not just technically fluent in Arabic but are also culturally aware and safe when deployed in the Middle East and North Africa.
Omar Abdelnasser, Fatemah Alharbi, Khaled Khasawneh, Ihsen Alouani, Mohammed E. Fouda
Compumacy for Artificial Intelligence Solutions · Computer Science Department, College of Computer Science and Engineering, Taibah University · Electrical and Computer Engineering Department, George Mason University · CSIT, Queen’s University Belfast
cs.CL, cs.AI
Submitted: 2026-02-03
Updated: 2026-09-29
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 87/100
The gist: SalamahBench introduces a unified benchmark for evaluating Arabic Language Models (ALMs) safety, addressing critical gaps in existing English-centric safety resources by providing category-aware
Key concepts
- MLCommons Hazard Taxonomy
- This is the standard set of 12 categories used to classify potential harms in AI models, such as hate speech or intellectual property violations. SalamahBench maps various Arabic datasets directly to these specific categories to ensure consistent, category-aware testing across different models.
- Majority Vote Safeguards
- Instead of relying on a single safety check, this method evaluates each model response using three different safeguard systems independently. A response is labeled unsafe only if the majority (at least two out of three) of these independent guards flag it, aiming for a more conservative and reliable safety assessment.
- Attack Success Rate (ASR)
- ASR measures how often an AI model generates a response that is classified as unsafe or harmful by the evaluation system. Researchers track both strict and loose versions of this metric to understand the frequency of harmful output across various safety settings.
Terminology
Summary
SalamahBench introduces a unified benchmark for evaluating Arabic Language Models (ALMs) safety, addressing critical gaps in existing English-centric safety resources by providing category-aware evaluation across 12 MLCommons hazard categories. This work is significant because it establishes a standardized, native-language framework necessary for rigorously assessing the robustness and alignment of ALMs to real-world Arabic linguistic and cultural contexts.
The gist
SalamahBench is a unified benchmark for evaluating the safety of ALMs, comprising 8,170 prompts across 12 different categories aligned with the MLCommons Safety Hazard Taxonomy.
Dataset Construction and Harmonization
The construction of SalamahBench follows a structured, multi-stage pipeline illustrated in Figure 2,
organized into three consecutive phases: preprocessing, AI filtering, and human validation. This process involves aggregating and harmonizing multiple existing datasets into a single corpus to ensure comprehensive coverage across all MLCommons hazard categories. The preprocessing phase normalizes raw datasets, filters them for quality and ambiguity using an AI-based judge (Claude Sonnet 4.5 and GPT-5), and concludes with a rigorous human verification pipeline involving a Human Harm Verifier and a Human Category Validator to ensure annotation consistency, correctness, and overall dataset quality.
Evaluation Protocol
The evaluation protocol adopts a two-stage process: first, each prompt from the benchmark is provided to the target LM to generate a response; second, this generated response is evaluated using a safeguard model. To derive a more conservative estimate of unsafe model behavior, the study introduces Majority Vote Safeguards,
where each generated response is independently evaluated by three distinct safeguard systems, and labeled as unsafe if it is flagged by a majority of the guards.
Safety Metrics and Analysis
The researchers employ three complementary safety metrics to assess ALMs:
-
Attack Success Rate (ASR): This measures the proportion of model responses classified as unsafe or containing harmful content, with both strict and loose variants reported.
-
Macro-Averaged ASR: This metric is used to ensure
equal treatment of all harm categories irrespective of their prompt frequency,
calculating an unweighted mean across all harm categories. -
Refusal Rate (RR): This measures the fraction of prompts for which the model generates an explicit refusal, detected using Qwen3Guard.
Model Performance and Findings
The empirical results reveal substantial variation in safety alignment
across models and safeguards. Key findings include:
** Fanar 2 consistently achieves the lowest attack success rates across all evaluation settings, including aggregate and macro-averaged category analyses.**
In contrast, Jais 2 exhibits the highest vulnerability, with ASR values reaching 24.2% under strict generative Qwen3Guard and remaining elevated across other safeguards. The analysis further demonstrates that safety alignment in LLMs is category-dependent,
as models with strong overall safety performance do not uniformly exhibit robustness across all harm domains, such as Intellectual Property and Sexual Content.
Finally, the study concludes that native ALMs are not reliable when used directly as safety guards,
showing substantially lower accuracy compared to dedicated safeguard systems.
Safeguard Model Validation
The effectiveness of specialized models is validated against human gold labels. Qwen3Guard achieved the highest accuracy in response-level safety classification at 90.8%, demonstrating strong alignment with human judgments.
The paper also evaluates native ALMs as self-guarding systems,
finding that even Fanar 2 achieves less than 50% accuracy against human annotations, reinforcing the necessity of specialized safeguard architectures.
Limitations and Future Directions
The study acknowledges several limitations, including the dataset's lack of full linguistic diversity across all Arabic dialects, an imbalanced distribution across certain high-risk categories, and the current inability of most safeguards to reliably distinguish between different response types (e.g., safe redirection versus refusal). Future work is directed toward developing native Arabic safeguard models that are explicitly trained for safety reasoning in Arabic contexts
and expanding the benchmark to support multi-turn interactions.
Category Mapping Transparency
A crucial component of the methodology is the detailed mapping of existing datasets to the MLCommons taxonomy. This transparency ensures consistency, as demonstrated by Appendix A, which documents how categories from sources like RTP-LX, AraSafe, and X-Safety are systematically mapped to the 12 hazard categories defined in Table 1. For instance, Insult
and Bias
from Microsoft RTP-LX are mapped to the Hate Speech
category. This rigorous alignment enables a unified and category-aware evaluation across datasets.
Prompt Design for Classification
The paper details the prompt templates used by safeguard models, such as Qwen3Guard, which use a modular design consisting of a shared system prefix enforcing strict classification behavior and a category-specific suffix providing the definition and decision criteria for each hazard category.
Improvements for AI systems
Here are the specific improvements that can be made to AI systems based on the provided research paper, along with what those improved systems could achieve:
) 1. Implement a Category-Aware Safety Evaluation Pipeline (SalamahBench):
The system should integrate a unified benchmark like SalamahBench, which harmonizes heterogeneous datasets across 12 MLCommons safety hazard categories. This moves evaluation beyond simple pass/fail metrics to fine-grained, category-level vulnerability mapping.
) Improved System Capability: The AI system will be able to detect and categorize specific types of harm with high precision (e.g., distinguishing between Violent Crimes,
Hate Speech,
and Indiscriminate Weapons
). This allows for targeted safety interventions rather than broad, generalized guardrails.
) 2. Develop Native-Language Safety Guard Models:
The system should prioritize training dedicated safeguard models specifically on Arabic linguistic nuances, cultural context, and dialectal variations (as suggested by the findings that native ALMs perform poorly as judges). This involves fine-tuning existing architectures (like Qwen3Guard or Llama Guard) using Arabic-specific safety data.
) Improved System Capability: The AI system will exhibit significantly higher accuracy in classifying nuanced Arabic harms, such as indirect harm, cultural framing of sensitive topics, and context-dependent harmful intent. It will be less prone to false negatives caused by translation artifacts compared to English-centric models.
) 3. Deploy Multi-Stage Safeguard Ensembles (Majority Vote):
Instead of relying on a single safety classifier, the system should employ an ensemble strategy where multiple distinct safeguard models (e.g., Qwen3Guard, Llama Guard 4, PolyGuard) evaluate the output and use a majority-vote aggregation rule to determine the final safety status.
) Improved System Capability: The system will achieve a significantly lower Attack Success Rate (ASR), particularly in high-risk domains like Violent Crimes
and Indiscriminate Weapons,
by requiring consensus across diverse safety models. This creates a more robust, conservative defense against adversarial prompts.
) 4. Enhance Safeguard Granularity and Disambiguation:
Integrate specialized disambiguation prompts (as seen in the Qwen3Guard methodology) when the initial coarse classification flags a response under an ambiguous category (like Sexual Content
). The system should be trained to recognize context that allows it to re-evaluate the prompt against more specific hazard definitions.
) Improved System Capability: The AI system will reduce mislabeling errors, correctly distinguishing between benign sexual content and actual Sex-Related Crimes
or Child Sexual Exploitation.
This minimizes false positives in sensitive areas while maintaining high sensitivity to genuine policy violations.
) 5. Transition from Static Prompts to Contextual Safety Evaluation:
Extend the evaluation methodology beyond single-turn prompts by incorporating multi-turn interaction analysis, allowing safeguard models (like FanarGuard) to assess safety based on the cumulative context of a conversation, not just isolated inputs or outputs.
) Improved System Capability: The AI system will be capable of identifying conversational drift
or gradual escalation of harm over several turns, enabling proactive intervention before a single turn crosses a hard threshold.
) 6. Implement Cross-Lingual Safety Consistency Checks:
Utilize multilingual datasets (like X-SAFETY and LinguaSafe) during the training and fine-tuning phases of safety classifiers to ensure that safety learned in one language (e.g., English) is consistently applied across other languages (Arabic).
) Improved System Capability: The system will demonstrate better cross-lingual generalization, reducing the performance gap observed between models aligned primarily for English and those operating natively in Arabic.
Sources
- Language Models are Few-Shot Learners
- Ethical and social risks of harm from Language Models
- A General Language Assistant as a Laboratory for Alignment
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Fanar: An Arabic-Centric Multimodal Generative AI Platform
- ALLaM: Large Language Models for Arabic and English
- AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
- Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Constitutional AI: Harmlessness from AI Feedback
- Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
- Red Teaming Language Models with Language Models
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- Qwen3Guard Technical Report
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering