🛰️ Daily AI Frontier
‹ back to 2026-09-11

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

Research Multimodal & Generative

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Representative image for OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

Merged summary

TL;DR - OmniHallu is a unified framework and benchmark for detecting hallucinations across multimodal comprehension and generation tasks involving images, video, audio, and text. It matters because prior detectors are often limited to a single modality or task type.

  • OmniHallu-Bench contains 10,000 samples with claim-level human annotations across six text-media conversion tasks.
  • A multi-agent architecture decomposes outputs into atomic claims, uses modality-specific experts for verification, and aggregates their evidence through structured reasoning.
  • A preference-optimized trainable verifier approximates the multi-agent decisions while reducing expert calls by 66% with minimal reported performance loss.
  • Experiments identify a consistent modality-dependent performance gradient and provide fine-grained analysis of cross-modal hallucination patterns.

Sources (1)

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

arXiv cs.CL Jianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo 2026-09-10 arXiv:2609.11244
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-22 14:22:53.760940 UTC

TL;DR - OmniHallu is a unified framework and benchmark for detecting hallucinations across multimodal comprehension and generation tasks involving images, video, audio, and text. It matters because prior detectors are often limited to a single modality or task type.

  • OmniHallu-Bench contains 10,000 samples with claim-level human annotations across six text-media conversion tasks.
  • A multi-agent architecture decomposes outputs into atomic claims, uses modality-specific experts for verification, and aggregates their evidence through structured reasoning.
  • A preference-optimized trainable verifier approximates the multi-agent decisions while reducing expert calls by 66% with minimal reported performance loss.
  • Experiments identify a consistent modality-dependent performance gradient and provide fine-grained analysis of cross-modal hallucination patterns.
item →