HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering
TL;DR - HalluTruthQA is an Arabic question-answering benchmark that evaluates not only hallucination detection, but also error localization, factual verification, and explanation. It provides richer supervision for measuring factual reliability in Arabic LLMs.
- Contains 2,400 expert-curated examples across Islamic knowledge, history, science, and geography.
- Includes verified answers, hallucination labels, candidate answers, character-level error spans, explanations, and hallucination types.
- Zero-shot evaluation of four open-source LLMs found that no model performed best across all four tasks.
- Best reported scores were 0.880 Macro-F1 for detection, 0.516 span F1 for localization, 0.852 LO-Score for verification, and 0.644 for explanation.