OmniHallu: Unified Hallucination Detection for Cross-Modal Comprehension and Generation in Multimodal Large Language Models
AuthorsJianjiang Yang, Peihang Li, Shanqing Xu, Mengchen Qian, Lu Zhang, Meng Luo
Resources
OmniHallu aims to detect hallucinations across image, video, and audio tasks using claim-level verification and a more efficient learned verifier.
Key results
Human-annotated samples covering six cross-modal tasks
Macro-F1 point improvement over the strongest baseline on a task
Claim-level hallucination detection score
Claim-level hallucination detection score
Reduction achieved by the GRPO-trained verifier filter
What the paper found
OmniHallu presents a unified claim-level hallucination detector for multimodal large language models across six bidirectional tasks: image-to-text, video-to-text, audio-to-text, text-to-image, text-to-video, and text-to-audio. Its OmniHallu-Bench contains 10,000 human-annotated samples spanning image, video, and audio inputs and outputs. The method decomposes captions or prompts into atomic claims with GPT-4.1, verifies them using modality-specific experts such as Grounding DINO, Qwen2.5-VL-72B, VideoLLaMA3, and Qwen2-Audio-7B-Instruct, then aggregates evidence with GPT-5.2. Compared with Self-Check and UNIHD, it improves claim-level macro-F1 by 3.4 to 8.1 points, reaching 82.95 on image-to-text, 77.14 on video-to-text, and 74.11 on audio-to-text; the largest gains occur in audio, where individual models are weaker. The benchmark also reveals a performance gradient from image to video to audio, and from comprehension to generation. Removing atomic claim decomposition reduces macro-F1 by up to 7.93 points, confirming that fine-grained localization is central. A GRPO-trained verifier initialized from Qwen2.5-VL-7B acts as a low-cost filter, reducing expert calls by 66 percent while preserving performance. The evaluation includes outputs from major systems including OpenAI’s GPT-4.1, Google’s Gemini-2.5-Pro, DALL-E 3, Stable Diffusion 3.5 Large, and Open-Sora, positioning OmniHallu as a cross-modal evaluation framework rather than a detector tied to one model family.
Original abstract
While Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse tasks, they suffer from hallucinations where generated outputs contradict or misrepresent input semantics. Existing research typically addresses hallucination detection within a single modality or task type, limiting generalizability. We introduce OmniHallu, a unified hallucination detection framework spanning both comprehension and generation tasks across image, video, and audio modalities. We contribute OmniHallu-Bench, a 10,000-sample benchmark with claim-level human annotations covering six cross-modal tasks: image-to-text (I2T), video-to-text (V2T), audio-to-text (A2T), text-to-image (T2I), text-to-video (T2V), and text-to-audio (T2A). Our multi-agent architecture decomposes model outputs into atomic claims, verifies them through modality-specific experts, and aggregates evidence via structured reasoning. We further propose a preference-optimized trainable verifier that approximates the multi-agent decision boundary, reducing expert calls by 66% with minimal performance loss. Extensive experiments reveal a consistent modality-dependent performance gradient and provide fine-grained insights into cross-modal hallucination patterns.
Read the original paperMore in Multimodal AI
Browse all 61 papers →Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation
Mohammed Irfan Kurpath, Jaseel Muhammad Kaithakkodan, Sahal Shaji Mullappilly, Ivan Laptev, Hisham Cholakkal
A compact embedding model unifies text, speech, audio, images, video, and documents in one search space without sacrificing the original text capabilities.
Qwen3.8-Omni: Towards Native Omni-Modal Agents
Qwen Team
Qwen3.8-Omni-Flash combines native text, audio, and video reasoning with long-context agentic planning and open-source tools for practical multimodal workflows.
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Ruibin Yuan, Jiahao Pan, Junyan Jiang, Zhiyue Wu, Ziya Zhou, Jiankai Sun, Yizhi Li, Ge Zhang, Yicheng Gu, Zeyue Tian, Junyu Dai, Hanfeng Lin, Kai Li, Shangda Wu, Xuanjie Liu, Jiaming Wang, Zihan Liu, Yue Wang, Yinghao Ma, Hanzhi Yin, Kangrui Chen, Xinyue Zhang, Ziyang Ma, Mengqi Liao, Hejia Zhao, Guowei Huang, Chao Yan, Lei Ke, Jianwei Yu, Bei Liu, Joe Guo, Liumeng Xue, Gus Xia, Wei Xue, Yike Guo
YuE2 turns readable musical scores into high-quality full songs, combining controllable composition with end-to-end audio generation.