🛰️ Daily AI Frontier
‹ back to 2026-07-23

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Research LLMs & Foundation Models

Ranking

Overall 77
Content 90
Popularity 46

Observed public metrics from 1 member.

Merged summary

TL;DR - HalluTruthQA is an Arabic question-answering benchmark that evaluates not only hallucination detection, but also error localization, factual verification, and explanation. It provides richer supervision for measuring factual reliability in Arabic LLMs.

  • Contains 2,400 expert-curated examples across Islamic knowledge, history, science, and geography.
  • Includes verified answers, hallucination labels, candidate answers, character-level error spans, explanations, and hallucination types.
  • Zero-shot evaluation of four open-source LLMs found that no model performed best across all four tasks.
  • Best reported scores were 0.880 Macro-F1 for detection, 0.516 span F1 for localization, 0.852 LO-Score for verification, and 0.644 for explanation.

Sources (1)

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

arXiv cs.CL Abdessalam Bouchekif, Mohammed-En-Nadhir Zighem, Salah Eddine Bekhouche, Hichem Telli, Somaya Eltanbouly, Shahd Gaben, Heba Sbahi, Samer Rashwani, Mutaz Al-Khatib, Emad Mohamed, Mohammed Ghaly, Abdenour Hadid 2026-07-22 arXiv:2607.20219
Public signals Hugging Face upvotes 0
Providers: Hugging Face · Upvotes 0 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-08-21 14:37:12.184068 UTC

TL;DR - HalluTruthQA is an Arabic question-answering benchmark that evaluates not only hallucination detection, but also error localization, factual verification, and explanation. It provides richer supervision for measuring factual reliability in Arabic LLMs.

  • Contains 2,400 expert-curated examples across Islamic knowledge, history, science, and geography.
  • Includes verified answers, hallucination labels, candidate answers, character-level error spans, explanations, and hallucination types.
  • Zero-shot evaluation of four open-source LLMs found that no model performed best across all four tasks.
  • Best reported scores were 0.880 Macro-F1 for detection, 0.516 span F1 for localization, 0.852 LO-Score for verification, and 0.644 for explanation.
item →