🛰️ Daily AI Frontier
‹ back to 2026-07-23

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Research LLMs & Foundation Models

Merged summary

TL;DR - HalluTruthQA is an Arabic question-answering benchmark that evaluates not only hallucination detection, but also error localization, factual verification, and explanation. It provides richer supervision for measuring factual reliability in Arabic LLMs.

  • Contains 2,400 expert-curated examples across Islamic knowledge, history, science, and geography.
  • Includes verified answers, hallucination labels, candidate answers, character-level error spans, explanations, and hallucination types.
  • Zero-shot evaluation of four open-source LLMs found that no model performed best across all four tasks.
  • Best reported scores were 0.880 Macro-F1 for detection, 0.516 span F1 for localization, 0.852 LO-Score for verification, and 0.644 for explanation.

Sources (1)

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

arXiv cs.CL Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem, Salah Eddine Bekhouche, Hichem Telli, Somaya Eltanbouly, Shahd Gaben, Heba Sbahi, Samer Rashwani, Mutaz Al-Khatib, Emad Mohamed, Mohammed Ghaly, Abdenour Hadid 2026-07-22 arXiv:2607.20219

TL;DR - HalluTruthQA is an Arabic question-answering benchmark that evaluates not only hallucination detection, but also error localization, factual verification, and explanation. It provides richer supervision for measuring factual reliability in Arabic LLMs.

  • Contains 2,400 expert-curated examples across Islamic knowledge, history, science, and geography.
  • Includes verified answers, hallucination labels, candidate answers, character-level error spans, explanations, and hallucination types.
  • Zero-shot evaluation of four open-source LLMs found that no model performed best across all four tasks.
  • Best reported scores were 0.880 Macro-F1 for detection, 0.516 span F1 for localization, 0.852 LO-Score for verification, and 0.644 for explanation.
item →