HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering".
Tom: Large language models can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify.
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Okay, so let's talk about who wrote this and what the title actually means. The paper is called "HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering".
Jane: It’s clear that they’re focusing specifically on Arabic question answering because they know how tricky it can be to evaluate these models accurately in that language. The authors are listed as Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem Salah Eddine Bekhouche, Hichem Telli, Somaya Eltanbouly, Shahd Gaben, Heba Sbahi, Samer Rashwani, Mutaz Al-Khatib and Emad Mohamed.
Lu: I see a lot of names here from various universities across the Middle East and beyond. It shows this is a collaborative effort bringing together expertise from different regions to tackle this specific linguistic challenge.
Meng: The title itself really sets expectations for what you’re getting—it’s not just about finding errors; it’s about detecting, localizing, and explaining them with great detail. That level of granularity suggests they're aiming for a much more precise evaluation tool than what we usually see in the field.
Lalam: I think the implication is that current benchmarks are too basic. They’ve built something specifically designed to probe the limits of Arabic LLMs in factual recall and reasoning, and that’s a big step forward for testing their actual capabilities.
The paper's summary: Tom: So what does this benchmark actually contain? Essentially, they’ve put together two thousand four hundred expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography <ref:2607.20219#pg0,2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge>. That covers a lot of ground.
Jane: Those two thousand four hundred examples are paired with a model's answer and a verified reference answer <ref:2607.20219#pg0>. And for the answers that are wrong—the hallucinated ones—they provide six candidate answers for verification. Plus they include character-level erroneous spans and human-written explanations for those errors.
Lu: The structure is pretty thorough, right? They aren't just giving you a label; they’re giving you the evidence needed to see the actual mistake in the text and understand what kind of mistake it is.
Meng: It sounds like they are tackling some really tricky issues, especially in those knowledge domains where things like temporal reasoning or source attribution can get mixed up. That complexity is exactly what we need to test for real-world applications.
Lalam: The key takeaway from the summary is that they are assessing where and why models fail, not just if they failed overall. They are looking at the nuance of the error itself, which should give us much better data on model weaknesses in Arabic QA.
The paper's improvements: Tom: Now, let’s talk about what this benchmark actually improves over what came before. The authors point out a few key enhancements in their approach to evaluation.
Jane: They introduce a multi-layered annotation scheme, which means they aren't stopping at just a response-level label. They are adding character-level spans and human explanations for hallucinated answers, which is important because they note that response-level labels can’t tell you the exact extent of the error.
Lu: They also have these four distinct evaluation tasks: binary detection, span localization, explanation evaluation using an LLM as a judge, and multiple-choice factual verification using an LO-Score metric. That gives them different ways to measure different types of failure.
Meng: It’s interesting how they use those candidate answers for verification; it forces the model to actually justify its choice against other plausible facts, which is a much stronger test than just checking if the answer matches one correct option.
Lalam: One big improvement they highlight is moving beyond simple binary detection to assessing where and why models fail and whether they can recover correct information, which really pushes the evaluation into a more practical zone for debugging.
Conclusion: Tom: So, wrapping things up on "HalluTruthQA," what’s the big picture here? Essentially, this benchmark gives us a much finer lens through which to look at Arabic QA models. It shows that simply having a model that can generate fluent Arabic isn't enough; we need tools to check the details.
Jane: The paper makes it clear that response-level labels alone are insufficient for deep error analysis, and HalluTruthQA provides the tools—the localization, the explanation evaluation—to see what’s actually going wrong inside the answer.
Lu: The results show that different models struggle with different things across these tasks; some are good at spotting hallucinations but bad at locating them precisely, while others do better on other aspects. That tells us where we should focus our future training efforts.
Meng: From a practical standpoint, the error analysis they did is really helpful because it groups errors into categories like temporal normalization or source attribution issues, which lets us target specific weaknesses in the model’s knowledge base for improvement.
Lalam: Ultimately, this work provides a much more granular way to measure Arabic LLM performance on factual tasks. It gives researchers concrete data points on error types that we can use to guide future development of more robust and trustworthy language models.
Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem Salah Eddine Bekhouche, Hichem Telli, Somaya Eltanbouly, Shahd Gaben, Heba Sbahi, Samer Rashwani, Mutaz Al-Khatib, Emad Mohamed, Mohammed Ghaly
Hamad Bin Khalifa University, Qatar · University of Biskra, Algeria · University of the Basque Country, Spain · Nazarbayev University, Kazakhstan · Universiti Malaysia Kelantan, Malaysia
cs.CL
Submitted: 2026-07-22
Updated: 2026-10-03
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 89/100
The gist: Large language models can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify.
Key concepts
- Binary Hallucination Detection
- This task checks if a model's answer is entirely correct or not. It uses Macro-F1 as the main metric because it treats both correct and incorrect answers equally, which is useful when you have an uneven number of true versus false answers.
- Span-level Localization
- This involves pinpointing the exact characters in a model's answer that are wrong. The performance is measured by how well the system can overlap its predicted faulty text span with the actual incorrect segment in the reference answer.
- Explanation Evaluation
- This assesses whether a human-written explanation for a hallucinated answer is helpful or accurate. An LLM acts as a judge, scoring explanations from 0 (completely wrong) to 2 (completely correct), helping to understand the nature of the model's failure.
- Multiple-Choice Factual Verification
- Instead of just checking if an answer is wrong, this task requires selecting the single best option from six candidate answers. The LO-Score metric evaluates both whether a hallucination exists and which candidate answer is factually correct.
Terminology
Summary
Large language models can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. This paper introduces HalluTruthQA, a fine-grained benchmark for hallucination evaluation in Arabic QA that moves beyond binary detection to assess where and why models fail and whether they can recover correct information.
How it works
The benchmark contains 2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography (Table 1). Each example pairs an Arabic question and a model-generated answer with a verified reference answer, a binary hallucination label, six candidate answers for factual verification, and character-level erroneous spans accompanied by human-written explanations for hallucinated answers. This multi-layer annotation scheme includes response-level hallucination labels, character-level erroneous spans, humanwritten explanations, and six candidate answers for factual verification
(Page 2).
Data Description and Annotation
The annotation process followed two passes: expert annotation followed by verification by trained research assistants (Page 3). Four domain experts annotated disjoint domain-specific subsets, ensuring each subset was annotated by an expert familiar with the corresponding knowledge area and its factual verification needs (Page 15). Annotators checked errors involving entities, dates, numerical values, unsupported claims, fabricated references, and incorrect attributions
(Page 3). The annotation process was based on meaning rather than exact string matching; for example, corresponding Hijri and Gregorian dates were accepted when the experts judged them to refer to the same historical event
(Page 15).
Evaluation Tasks and Metrics
The benchmark supports four evaluation tasks:
-
Binary Hallucination Detection: Predicting whether a response is hallucinated or non-hallucinated, using
Macro-F1 as the primary detection metric, as it equally weights both classes and is robust to label imbalance
(Page 6). -
Span-level Localization: Localizing erroneous spans at the character level, where performance is based on
overlap between predicted and gold erroneous spans
(Page 6). -
Explanation Evaluation: Using an LLM-as-a-judge protocol, where the judge assigns a score from 0 to 2 (0: incorrect/irrelevant; 1: partially correct; 2: complete/-correct), normalized to [0, 1] (Page 5).
-
Multiple-Choice Factual Verification: Selecting the correct option from six candidate answers, using the LO-Score metric which
evaluates hallucination detection and option selection
(Page 6).
Model Evaluation and Findings
The evaluation involved four open-source LLMs: ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, and SILMA in a zero-shot setting (Page 5). Results show that no single model achieves the strongest performance across all tasks
(Page 6). For instance, ALLaM7B shows the strongest detection performance, achieving the highest global Macro-F1 score (0.880)
(Page 6), while Qwen3-32B achieves the best global localization performance, with an F1-Sp of 0.516
(Page 6). Domain-specific insights reveal that Islamic Knowledge is affected by evidencegrounding and source-attribution errors, whereas history, science, and geography are mainly affected by factual contradictions and fabricated details
(Page 7).
Error Analysis
The error analysis groups the most frequent model-level errors into four main families:
-
Temporal, Numeric, and Unit Normalization (39.0%):
These errors occur when models fail to recognize that two different formats express the same fact
(Page 8). -
Source Attribution and Evidence Verification (22.2%):
These errors are especially common in the Islamic domain
(Page 8). -
Correct Short Answer with Hallucinated Support (17.1%):
Detection models often miss these cases because they rely too much on the overlap with the correct answer and ignore the hallucinated information in the explanation
(Page 8). -
Fully Incorrect Answer (11.3%):
These errors occur when the model’s main answer is factually incorrect
(Page 8).
REFERENCES
Samir Abdaljalil, Hasan Kurban, and Erchin Serpedin. 2025. Halluverse25: Fine-grained multilingual benchmark dataset for llm hallucinations. arXiv preprint arXiv:2503.07833.
Samir Abdaljalil, Parichit Sharma, Erchin Serpedin, and Hasan Kurban. 2026. Halluversem 3: A multitask multilingual benchmark for hallucination in llms. arXiv preprint arXiv:2602.06920.
Ali Abdelaal, Mohammed Nader Al Haffar, Mahmoud Fawzi, and Walid Magdy. 2026. Islamicmmlu: a benchmark for evaluating llms on islamic knowledge. arXiv preprint arXiv:2603.23750.
Aisha Alansari and Hamzah Luqman. 2025. AraHalluEval: A fine-grained hallucination evaluation framework for Arabic LLMs. In Proceedings of The Third Arabic Natural Language Processing Conference, pages 148–161, Suzhou, China. Association for Computational Linguistics.
Aisha Alansari and Hamzah Luqman. 2026a. Halluscore: Large language model hallucination question answering benchmark. arXiv preprint arXiv:2605.17007.
Aisha Alansari and Hamzah Luqman. 2026b. Large language models hallucination: A comprehensive survey. arXiv preprint arXiv:2510.06265.
Fakhraddin Alwajih, Abdellah El Mekki, Hamdy Mubarak, Majd Hawasly, Abubakr Mohamed, and Muhammad Abdul-Mageed. 2025. Palmx 2025: The first shared task on benchmarking llms on arabic and islamic culture. In Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pages 774–789.
Abdessalam Bouchekif, Somaya Eltanbouly, Samer Rashwani, Shahd Gaben, Mutaz AlKhatib, Heba Sbahi, Emad Mohamed, and Mohammed Ghaly. 2026a. Qias 2026: Overview of the shared task on islamic inheritance reasoning.
Abdessalam Bouchekif, Samer Rashwani, Emad Soliman Ali Mohamed, Mutaz Alkhatib, Heba Sbahi, Shahd Gaben, Wajdi Zaghouani, Aiman Erbad, and Mohammed Ghaly. 2025a. QIAS 2025: Overview of the shared task on islamic inheritance reasoning and knowledge assessment. In Proceedings of The Third Arabic Natural Language Processing Conference: Shared Tasks, pages 851–860, Suzhou, China.
Abdessalam Bouchekif, Samer Rashwani, Heba Sbahi, Shahd Gaben, Mutaz Al Khatib, and Mohammed Ghaly. 2025b. Assessing large language models on islamic legal reasoning: Evidence from inheritance law evaluation. In Proceedings of The Third Arabic Natural Language Processing Conference, pages 246–257, Suzhou, China.
Shiqi Chen, Yiran Zhao, Jinghan Zhang, I-Chun Chern, Siyang Gao, Pengfei Liu, and Junxian He. 2023. FELM: Benchmarking factuality evaluation of large language models. In Thirty-seventh Conference on Neural Information Processing Systems Datasets and Benchmarks Track.
Passant Elchafei and Mervat Abu-Elkheir. 2025. Hallucination detectives at SemEval-2025 task 3: Span-level hallucination detection for LLMgenerated answers. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), pages 601–606, Vienna, Austria.
Fanar Team, Ummar Abbas, Mohammad Shahmeer Ahmad, Firoj Alam, Enes Altinisik, Ehsannedin Asgari, Yazan Boshmaf, Sabri Boughorbel, Sanjay Chawla, Shammur Chowdhury, Fahim Dalvi, Kareem Darwish, Nadir Durrani, Mohamed Elfeky', Ahmed Elmagarmid', Mohamed Eltabakh', Masoomali Fatehkia', Anastasios Fragkopoulos', Maram Hasanain, and 23 others. 2025. Fanar: An arabic-centric multimodal generative AI platform. Preprint, arXiv:2501.13944.
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin', and 1 others. 2025a.
Improvements for AI systems
-
A multi-layered evaluation pipeline can be implemented for Arabic QA systems that moves beyond binary correctness to granular error analysis. This system would perform
response-level hallucination detection, exact character-level span localization, human-written explanations, and candidate-based factual verification.
-
The improved system can localize errors precisely using character offsets by calculating
partialcredit span-level F1 (F1-Sp) following (Sky et al., 2024),
addressing the limitation thatresponse-level labels cannot identify the precise nature or extent of an error.
-
The system can provide actionable feedback on why an answer is wrong by leveraging a structured prompt for explanation evaluation, scoring models based on
Error Identification (0–1)
andFactual Correction (0–1)
dimensions to generate justifications. -
Model training can be guided by domain-specific error analysis; for instance, the system can be trained to recognize that in the Islamic Knowledge domain, errors often involve
evidence grounding and source-attribution errors.
-
The system can improve factual reasoning by incorporating checks for specific error families identified in the taxonomy, such as recognizing
Temporal, Numeric, and Unit Normalization
errors (e.g., distinguishing between different date formats like40 AH
vs.661 CE
).
Abstract
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer. We introduce HalluTruthQA, a fine-grained benchmark for hallucination evaluation in Arabic question answering. The benchmark contains 2,400 expert-curated examples across four knowledge-intensive domains: Islamic knowledge, history, science, and geography. Each example pairs an Arabic question and a model-generated answer with a verified reference answer, a binary hallucination label, and six candidate answers for factual verification. Hallucinated answers additionally include character-level erroneous spans, human-written explanations, and macro- and micro-level hallucination types. We evaluate four open-source LLMs, ALLaM-7B, Falcon-H1R-7B, Qwen3-32B, and SILMA, in a zero-shot setting across hallucination detection, span-level localization, factual verification, and explanation evaluation. Results show that these tasks capture different abilities: no single model performs best across all tasks. The best scores are 0.880 Macro-F1 for detection, 0.516 F1-Sp for localization, 0.852 LO-Score for factual verification, and 0.644 for explanation evaluation. These findings show that hallucination evaluation should move beyond response-level detection toward the localization, verification, and explanation of factual errors.
Sources
- HalluVerse25: Fine-grained Multilingual Benchmark Dataset for LLM Hallucinations
- Halluverse-M^3: A multitask multilingual benchmark for hallucination in LLMs
- IslamicMMLU: A Benchmark for Evaluating LLMs on Islamic Knowledge
- HalluScore: Large Language Model Hallucination Question Answering Benchmark
- Large Language Models Hallucination: A Comprehensive Survey
- QIAS 2026: Overview of the Shared Task on Islamic Inheritance Reasoning
- MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
- Fanar: An Arabic-Centric Multimodal Generative AI Platform
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering