Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection
King Fahd University of Petroleum and Minerals · SDAIA-KFUPM Joint Research Center for AI · University of Tabuk
cs.CL
Submitted: 2026-08-10
Updated: 2026-10-02
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: Mawqif-v2 is an extension of the original Mawqif dataset, designed as a benchmark for cross-target stance detection in Arabic.
Terminology
Summary
Mawqif-v2 is an extension of the original Mawqif dataset, designed as a benchmark for cross-target stance detection in Arabic. It consists of 996 manually annotated Arabic tweets collected from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is annotated with stance, sentiment, and sarcasm labels following the original Mawqif annotation scheme. The extension is intended as a held-out evaluation set for assessing model generalization to both semantically related and previously unseen targets, while the original Mawqif dataset is used for training and development.
The dataset construction involved target-specific search strategies for each target. For E-Cars, 5,006 tweets were collected from X (formerly Twitter) using the Twitter API with keyword-based search covering Arabic spelling variations of electric car
and electric cars,
spanning January 1, 2022 to February 15, 2022. For Trimester System, 3,000 tweets were collected using keywords related to three semesters,
the third academic term,
and similar expressions, spanning January 20, 2023 to February 8, 2023. For Women Driving, 8,450 tweets were collected using Octoparse Twitter Scraper with hashtags and lexical variations related to women driving, spanning January 2, 2016 to November 27, 2025.
After automatic filtering (removing non-Arabic tweets, duplicates, news/media accounts, advertisements, and adult-related content) and manual quality review, the final dataset contained 352 tweets for Women Driving, 332 tweets for E-Cars, and 312 tweets for Trimester System. Each tweet was independently annotated by three annotators, with a fourth annotator resolving disagreements. Inter-annotator agreement measured using Fleiss' Kappa was 0.564 for stance, 0.485 for sentiment, and 0.526 for sarcasm, indicating moderate agreement. Full agreement among all three annotators was observed in 62.95% of stance annotations, 54.22% of sentiment annotations, and 80.42% of sarcasm annotations.
The dataset includes two evaluation protocols. Protocol 1 evaluates related-target generalization: models are trained on COVID-19 Vaccine, Digital Transformation, and Women Empowerment targets, then evaluated on Women Driving, which is semantically related to Women Empowerment. Protocol 2 evaluates unseen-target generalization: models are trained on COVID-19 Vaccine and Digital Transformation, validated on Women Empowerment, and evaluated on E-Cars and Trimester System, which are from different domains.
Baseline experiments were conducted using four Arabic pretrained language models (AraBERT-v02, AraBERT-twitter, CAMeLBERT-da, MARBERT), three multilingual models (mBERT, DistilBERT, XLMRoBERTa), and three large language models in zero-shot settings (LLaMA 3.3 70B, Qwen 2.5 72B, JAIS 70B). The primary evaluation metric was Favg2, the macro-average F1 score over the Favor and Against classes.
Under Protocol 1, Qwen 2.5 achieved the best performance on the Women Driving test set (Favg2 = 73.66), while AraBERT-v02 obtained the best result among encoder-based models (Favg2 = 65.92). Under Protocol 2, Qwen 2.5 achieved the best overall test score (Favg2 = 74.75), with the highest result on E-Cars (Favg2 = 80.60), while MARBERT achieved the highest performance on the Trimester System target (Favg2 = 74.11). Among encoder-based models, MARBERT obtained the highest overall test score (Favg2 = 72.21).
The results show that no single model achieved the highest performance across all targets, highlighting differences across evaluation targets. The dataset supports research on Arabic stance detection, sentiment analysis, sarcasm detection, and multi-task learning. Limitations include the limited range of domains covered, the specific transfer scenarios evaluated, and the fact that data was collected only from X (formerly Twitter).
Improvements for AI systems
Improvements to AI systems:
-
Cross-target stance detection with semantic transfer learning: Train models on a diverse set of stance targets (e.g., COVID-19, digital transformation, women’s empowerment) and fine-tune them to generalize to semantically related but unseen targets (e.g., women driving). This improves the system’s ability to transfer stance knowledge across topics that share underlying social or political dimensions, reducing the need for target-specific labeled data.
-
Domain-agnostic stance detection via multi-task learning: Jointly train on stance, sentiment, and sarcasm labels (as provided in Mawqif-v2) to create a multi-task model that leverages shared linguistic cues (e.g., sarcasm often signals negative stance, sentiment polarity can disambiguate stance). This yields a more robust system that performs better on unseen targets by using auxiliary signals to compensate for missing target-specific context.
-
Zero-shot and few-shot stance detection for low-resource Arabic dialects: Use the benchmark to evaluate and fine-tune large language models (e.g., Qwen 2.5, LLaMA) for zero-shot stance detection on new Arabic targets. The improved system can then classify stance on entirely novel topics (e.g., new policy debates) without any labeled examples, using only the model’s pretrained knowledge and the evaluation protocols from the paper.
-
Target-aware representation learning: Develop a model that explicitly encodes the target’s semantic category (e.g., “transportation,” “education,” “gender rights”) using embeddings or attention mechanisms. This allows the system to adapt its stance prediction based on the target’s domain, improving performance on related targets (Protocol 1) and reducing confusion when targets share vocabulary but differ in stance context.
-
Sarcasm-aware stance calibration: Integrate sarcasm detection as a pre-processing or attention-modulating component. Since sarcasm often inverts literal sentiment, the improved system can automatically down-weight or flip sentiment signals when sarcasm is detected, leading to more accurate stance predictions on Arabic tweets where sarcasm is prevalent (as seen in the dataset’s sarcasm annotations).
-
Ensemble and model selection for target-specific robustness: Build an ensemble system that dynamically selects the best-performing model per target based on validation performance (e.g., MARBERT for Trimester System, Qwen for E-Cars). The improved system can route incoming tweets to the most suitable model, achieving higher overall accuracy across diverse targets than any single model.
-
Cross-lingual and cross-domain stance generalization: Use the multilingual models (mBERT, XLMRoBERTa) fine-tuned on Mawqif-v2 to transfer stance detection capabilities to other Arabic dialects or even other languages with similar social debates. The improved system can then be deployed in multi-country social media monitoring without retraining on each new language or dialect.
What the improved AI system can do:
-
Automatically detect public stance (favor/against/neutral) on new, previously unseen Arabic social media topics with high accuracy, even when no labeled data exists for that topic.
-
Distinguish between literal and sarcastic expressions of opinion, avoiding misclassification of sarcastic tweets as neutral or opposite stance.
-
Provide explainable stance predictions by leveraging sentiment and sarcasm signals jointly, helping analysts understand why a tweet was classified as favoring or opposing a target.
-
Adapt to new domains (e.g., environmental policy, labor rights) by using the transfer protocols from the paper, requiring only a small validation set to calibrate performance.
-
Serve as a robust benchmark-driven component in social listening tools, enabling real-time tracking of public opinion shifts across multiple Arabic-speaking regions and topics.
Sources
- ARBERT & MARBERT: Deep Bidirectional Transformers for Arabic
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- LLaMA: Open and Efficient Foundation Language Models
- Jais and Jais-chat: Arabic-Centric Foundation and Instruction-Tuned Open Generative Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering