ARAFA: An LLM-Generated Arabic Fact-Checking Dataset
cs.CL, cs.IR
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/chriskhalil/ARAFA
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources.
Terminology
Abstract
Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources. In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs). The dataset was constructed through a three-step pipeline: (1) claim generation from Arabic Wikipedia pages with supporting textual evidence, (2) claim mutation to generate challenging counterfactual claims with refuting evidence, and (3) an automatic validation step to validate that the generated claims are either supported or refuted by their accompanying evidence, or if the evidence does not provide enough information to judge the validity of the claims. The resulting dataset comprises 181,976 claim-evidence pairs labeled as supported, refuted, or not enough information. Human evaluation carried out on a test sample from the dataset demonstrated strong inter-annotator agreement (kappa = 0.89) using Cohen's Kappa for supported claims and (kappa = 0.94) for refuted claims. Automatic validation based on a human-evaluated sample achieved 86% accuracy for supported claims and 88% for refuted ones. To showcase Arafa's value as a resource for automatic Arabic fact-checking, four open-source transformer-based models were fine-tuned using Arafa, with the top-performing model achieving a Macro F1-score of 77% on the test data. In addition to Arafa being the first large-scale dataset for Arabic fact-checking, our framework presents a scalable approach for developing similar resources for other low-resource languages.
Sources
- AraStance: A Multi-Country and Multi-Domain Dataset of Arabic Stance Detection for Fact Checking
- Ta'keed: The First Generative Fact-Checking System for Arabic Claims
- FEVEROUS: Fact Extraction and VERification Over Unstructured and Structured information
- AraBERT: Transformer-based Model for Arabic Language Understanding
- Why do language models perform worse for morphologically complex languages?
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- CHEF: A Pilot Chinese Dataset for Evidence-Based Fact-Checking
- Adversarial Attacks and Defense on Texts: A Survey
- Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
- Understanding the effects of language-specific class imbalance in multilingual fine-tuning
- Context Length Alone Hurts LLM Performance Despite Perfect Retrieval
- Language Models as Fact Checkers?
- Explainable Automated Fact-Checking for Public Health Claims
- Assessing BERT's Syntactic Abilities
- The Llama 3 Herd of Models
- Fine-grained Fact Verification with Kernel Graph Attention Network
- A Richly Annotated Corpus for Different Tasks in Automated Fact-Checking
- Decoupled Weight Decay Regularization
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- A Survey of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering