NTH

OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models

AuthorsJianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo

September 12, 2026 2 min read
Watch on YouTube
The one-line take

OmniHallu aims to detect hallucinations across image, video, and audio tasks using claim-level verification and a more efficient learned verifier.

Key results

10,000
OmniHallu-Bench size

Human-annotated samples covering six cross-modal tasks

8.1
Maximum baseline improvement

Macro-F1 point improvement over the strongest baseline on a task

82.95
Image-to-text Macro-F1

Claim-level hallucination detection score

74.11
Audio-to-text Macro-F1

Claim-level hallucination detection score

66%
Expert-call reduction

Reduction achieved by the GRPO-trained verifier filter

What the paper found

OmniHallu presents a unified claim-level hallucination detector for multimodal large language models across six bidirectional tasks: image-to-text, video-to-text, audio-to-text, text-to-image, text-to-video, and text-to-audio. Its OmniHallu-Bench contains 10,000 human-annotated samples spanning image, video, and audio inputs and outputs. The method decomposes captions or prompts into atomic claims with GPT-4.1, verifies them using modality-specific experts such as Grounding DINO, Qwen2.5-VL-72B, VideoLLaMA3, and Qwen2-Audio-7B-Instruct, then aggregates evidence with GPT-5.2. Compared with Self-Check and UNIHD, it improves claim-level macro-F1 by 3.4 to 8.1 points, reaching 82.95 on image-to-text, 77.14 on video-to-text, and 74.11 on audio-to-text; the largest gains occur in audio, where individual models are weaker. The benchmark also reveals a performance gradient from image to video to audio, and from comprehension to generation. Removing atomic claim decomposition reduces macro-F1 by up to 7.93 points, confirming that fine-grained localization is central. A GRPO-trained verifier initialized from Qwen2.5-VL-7B acts as a low-cost filter, reducing expert calls by 66 percent while preserving performance. The evaluation includes outputs from major systems including OpenAI’s GPT-4.1, Google’s Gemini-2.5-Pro, DALL-E 3, Stable Diffusion 3.5 Large, and Open-Sora, positioning OmniHallu as a cross-modal evaluation framework rather than a detector tied to one model family.

Original abstract

While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our multi-agent architecture decomposes model outputs into atomic claims, verifies them through modality-specific experts, and aggregates evidence via structured reasoning. We further propose a preference-optimized trainable verifier that approximates the multi-agent decision boundary, reducing expert calls by 66% with minimal performance loss. Extensive experiments reveal a consistent modality-dependent performance gradient and provide fine-grained insights into cross-modal hallucination patterns.

Read the original paper

More in Multimodal AI

Browse all 61 papers →
02Multimodal

Qwen3.8-Omni: Towards Native Omni-Modal Agents

Qwen Team

Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.

Read analysis
03Multimodal

YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality

Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo

YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.

Read analysis