NTH

O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning

AuthorsMei Yuan, Qi Long, Qifeng Wu, Zhenyang Li, Yizhou Zhao, Lei Wang, Yang Liu, Min Xu

August 2, 2026 2 min read
Watch on YouTube
The one-line take

O-VAD uses object tracking and temporal reasoning to help AI inspectors explain anomalies in complex industrial video without task-specific training.

Key results

0.584
Phys-AD video AUROC

O-VAD's average video-level AUROC, outperforming GPT-5 at 0.503 and Qwen3-VL-32B at 0.513.

0.692
LiquidAD video AUROC

Training-free video-level AUROC on the multi-pipette liquid-transfer benchmark.

0.565
IPAD video AUROC

Average video-level AUROC on the periodic industrial-process benchmark.

0
State-tracking ablation failure rate

Precision, recall, and F1 fall to 0 on 3 of 4 Phys-AD subsets when state tracking is removed.

What the paper found

Researchers at Carnegie Mellon University introduce O-VAD, a training-free industrial video anomaly detector that replaces whole-frame classification with object-centric state evolution. Its three-stage pipeline uses GPT-5 from OpenAI to discover objects, SAM3 for concept-guided masks, CropFormer and SAM2 to build and recover spatiotemporal tubelets, and VLM queries to describe open-ended changes such as deformation, leakage, surface damage, and material release. A six-step reasoning chain then compares observed trajectories with expected process physics, assigns anomaly types and severity, localizes affected objects and frames, and applies visual verification to reduce false positives. Across Phys-AD, LiquidAD, and IPAD, O-VAD reaches a video-level AUROC of 0.584 on Phys-AD, exceeding GPT-5 at 0.503 and Qwen3-VL-32B at 0.513 without fine-tuning, domain knowledge, or predefined taxonomies. It also achieves 0.692 video-level AUROC on LiquidAD and 0.565 on IPAD, where anomalies involve multiple pipettes, periodic factory operations, and reference-based deviations. The ablation results identify state tracking as essential: removing it drives precision, recall, and F1 to 0 on 3 of 4 Phys-AD subsets. O-VAD produces interpretable reports that connect object states to causal diagnoses, although its multi-stage design adds latency and it remains weak on defects defined by invisible specifications, such as magnetic strength or exact actuation limits.

Original abstract

Industrial Video Anomaly Detection (IVAD) aims to identify anomalous objects and events in an industrial process, which is crucial for modern manufacturing and quality control systems. Existing VLM-based anomaly reasoning methods are capable of detecting open-ended anomalies in general domains. However, their performance declines in industrial settings characterized by intricate object transformations, strict physics, and procedural constraints. To tackle the complexity of such interaction-intensive detection, we introduce a training-free agentic framework for anomaly detection free of domain-specific knowledge, emphasizing object state evolution like humans inspectors. It is designed to track spatial-temporal dynamics and underlying transformations of detected objects over time, and then reason over the object-wise temporal state trajectories to identify abnormal objects in grounded frames. Our method overcomes limitations of prior approaches that rely on retraining on normal clips or injecting domain knowledge as context for test-time inference. Extensive experiments on three IVAD datasets demonstrate that our method outperforms frontier VLMs, agentic frameworks, and traditional VAD methods fine-tuned on the respective datasets, while providing interpretable reports over anomaly processes and types.

Read the original paper

More in Computer Vision

Browse all 58 papers →
02Cv

DyRAD: Radar Novel View Synthesis for Dynamic Driving Scenes

Merav Keidar, Tomer Borreda, Rajalakshmi Nandakumar, Or Litany

DyRAD builds moving radar views of driving scenes by combining tracked object motion with the radar’s physics, enabling more realistic and transferable autonomy testing.

Read analysis
03Cv

OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

Jiaxin Ge, Yiming Qin, Ji Xie, Haozhe Jiang, Xiaochuang Han, Junyi Zhang, Andrew Dai, Yinfei Yang, Jitendra Malik, Ranjay Krishna, Sewon Min, Haiwen Feng, Le Xue, Baifeng Shi, Trevor Darrell, XuDong Wang

This work maps when training models to generate images can make them better at understanding images, revealing both intuitive and surprising task-to-task benefits.

Read analysis