NTH

ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization

AuthorsYuxiong Xu, Kaiqing Lin, Bin Li, Haodong Li, Sheng Li

August 9, 2026 2 min read
Watch on YouTube
The one-line take

ThinkOmni teaches an audio-capable language model to explain why a recording is fake and pinpoint exactly when the manipulation occurs.

Key results

100K
FACoT dataset size

Forensic-Aware Chain-of-Thought training samples aggregated from eight public datasets.

80.74%
Cross-dataset accuracy

Mean accuracy on the ADD and Speech-Forensics evaluation sets.

85.15%
Cross-dataset F1

Mean F1 score on unseen ADD and Speech-Forensics data.

74.67%
Cross-dataset localization mAP

Temporal localization mean average precision across unseen datasets.

What the paper found

ThinkOmni reframes audio forgery detection and localization as evidence-grounded reasoning rather than implicit classification. Built on Qwen2.5-Omni, it combines semantic speech features, Wav2Vec2 XLSR-300M acoustic representations, and spectrogram vision through Forensic-Aware Modality-Incremental Learning, while its SAFE module fuses local semantic-acoustic interactions with global forgery context. The model is supervised with FACoT, a 100K-sample dataset assembled from eight public benchmarks, where 6.2K expert-reviewed seed samples guide Qwen3-Omni and Gemini-3-Pro-assisted chain-of-thought expansion; CLAP filtering removes weakly audio-grounded rationale components. Forensic-Consistent Multi-task Loss downweights reasoning-token dominance and jointly optimizes authenticity classification with timestamp-based boundary regression. On unseen ADD and Speech-Forensics data, ThinkOmni achieves 80.74% accuracy, 85.15% F1, and 74.67% temporal mAP, outperforming Qwen2-Audio and Qwen2.5-Omni baselines, while explicitly explaining cues such as prosodic mismatch, spectral artifacts, speaker inconsistency, and manipulation boundaries. The results position structured multimodal reasoning as a route to stronger cross-dataset generalization, although qualitative failures show that environmental recording effects can still cause false positives, missed forgeries, and overestimated boundaries.

Original abstract

Existing audio forgery detection and localization (AFDL) methods often overfit dataset-specific low-level artifacts, limiting their generalization to subtle, localized, and unseen manipulations. Recent audio large language model (ALLM)-based approaches cast AFDL as question answering but still model forensic evidence implicitly, without linking manipulation cues to predictions. To bridge this gap, we propose ThinkOmni, a reasoning-driven omni-modal large language model that jointly performs explicit forensic reasoning, spoofing detection, and temporal manipulation localization. To enable explicit reasoning supervision, we construct Forensic-Aware Chain-of-Thought (FACoT), a 100K-sample dataset with structured forensic evidence and reasoning annotations. Leveraging FACoT, we introduce Forensic-Aware Modality-Incremental Learning (FMIL), which progressively aligns semantic, acoustic, and spectral-visual representations with the LLM backbone to capture complementary forensic cues. We further propose Forensic-Consistent Multi-task Loss (FCML), which combines weighted cross-entropy with an adaptive localization loss to coordinate reasoning generation, spoofing detection, and temporal localization. Extensive experiments show that ThinkOmni achieves strong cross-dataset generalization in both detection and localization. Code, models, data, and inference examples are available at https://beyond0814.github.io/ThinkOmni/.

Read the original paper

More in Speech AI

Browse all 27 papers →
01Speech

Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training

Yifan Yang, Xiaoyu Yang, Zengrui Jin, Xian Shi, Yuxuan Wang, Yu Xi, Ziyang Ma, Qi Chen, Ruiyang Xu, Hui Wang, Dongchao Yang, Jin Xu, Xie Chen

A new CTC training strategy makes large-vocabulary LLM speech recognition far more memory-efficient while retaining competitive accuracy and fast streaming inference.

Read analysis