NTH

Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection

AuthorsXuechao Zou, Shun Zhang, Kai Li, Yi Zhou, Xinyu Sun, Yuhui Chen, Zhe Wu, Congyan Lang, Junliang Xing

August 13, 2026 2 min read
Watch on YouTube
The one-line take

A team of specialized AI investigators examines video texture, lighting, motion, and physics to expose deepfakes and explain its verdict.

Key results

100,000
FaceVid-Forensics-100K videos

Dataset size spanning 33 synthesis methods.

33
Synthesis methods

Forgery-generation methods covered by FaceVid-Forensics-100K.

69.87%
OOD accuracy

Accuracy of the full multi-agent system on unseen generators.

53.28%
OOD F1

Full-system F1 score on the out-of-domain test set.

5.83%
F1 improvement

Absolute gain over the strongest single-model baseline.

What the paper found

The paper introduces FaceVid-Forensics-100K, a deepfake benchmark containing 100,000 videos generated by 33 synthesis methods, including Seedance 2.0 and unseen systems such as OpenAI’s Sora 2. Unlike binary-only datasets, it provides verdict-consistent forensic explanations across texture, lighting, motion, and physics. Its detector uses four specialized agents based on Qwen2.5-VL-7B, with each agent independently reporting evidence; a judge agent, trained with supervised fine-tuning and GRPO, reconciles those reports and can additionally inspect sampled video frames. Training labels are aggregated with DeepSeek-V4 Pro from GPT-4o, Gemini 3.5-Flash, Qwen2.5-VL, Skyra, and VideoVeritas. On an out-of-domain split containing 20 unseen generators, the full system reaches 69.87% accuracy, 81.82% recall, and 53.28% F1, exceeding individual MLLMs such as GPT-4o and Gemini models. Its F1 rises from 47.45% for the strongest single-model baseline to 53.28%, an absolute improvement of 5.83 percentage points. The results indicate that independent evidence collection and explicit judge-based reconciliation generalize better than direct prediction, chain-of-thought, or multi-turn prompting, while preserving interpretable rationales.

Original abstract

The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deepfake video benchmarks provide limited coverage of recent synthesis methods and generally lack reliable fine-grained textual annotations. Meanwhile, conventional detectors and multimodal large language models (MLLMs), whether operating as a single model or relying on a single analytical perspective, often fail to capture subtle forgery artifacts, limiting their generalization to emerging AI-generated methods. To address these limitations, we introduce FaceVid-Forensics-100K, a large-scale deepfake video dataset comprising 100,000 videos and spanning 33 synthesis methods across face swapping, face reenactment, and entire-face synthesis, including recent generators such as Seedance 2.0. The dataset provides fine-grained textual annotations of visual observations and verdict-consistent forensic explanations, automatically synthesized through a multi-model aggregation and conflict-resolution pipeline powered by advanced MLLMs. Building on this benchmark, we propose a multi-agent forensic reasoning framework that employs four specialized domain-expert agents to independently analyze forgery cues from four perspectives: texture, lighting, motion, and physics. A judge agent then reconciles their reports to produce a final prediction together with an explanation. Extensive evaluations on out-of-domain test sets show that, despite being composed entirely of small open-source MLLMs, our framework outperforms all methods including closed-source GPT and Gemini models and ranks first across all reported metrics on this benchmark. The project page is available at https://xavierjiezou.github.io/ARGUS/.

Read the original paper

More in AI Safety

Browse all 39 papers →
02Safety

Language Models Are "Insecure" Reporters

Jenny Y. Huang, Jiameng Fan, Ahmed Imtiaz Humayun, Maximillian Chen, Tian Qin, Run Chen, Vidhya Navalpakkam, Hongxiang Gu

The study finds that language models often hide flaws that undermine their success stories, but a simple honesty instruction can make their reports dramatically more transparent.

Read analysis