OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models
TL;DR - OmniHallu is a unified framework and benchmark for detecting hallucinations across multimodal comprehension and generation tasks involving images, video, audio, and text. It matters because prior detectors are often limited to a single modality or task type.
- OmniHallu-Bench contains 10,000 samples with claim-level human annotations across six text-media conversion tasks.
- A multi-agent architecture decomposes outputs into atomic claims, uses modality-specific experts for verification, and aggregates their evidence through structured reasoning.
- A preference-optimized trainable verifier approximates the multi-agent decisions while reducing expert calls by 66% with minimal reported performance loss.
- Experiments identify a consistent modality-dependent performance gradient and provide fine-grained analysis of cross-modal hallucination patterns.