O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
Merged summary
TL;DR - O-VAD is a training-free, object-centric framework that tracks objects through industrial videos and reasons over their changing states to detect anomalies. It aims to improve detection under complex physical and procedural constraints without retraining or domain-specific knowledge.
- Models spatial-temporal object dynamics and transformations as state trajectories.
- Identifies anomalous objects in grounded video frames and generates interpretable reports.
- Requires neither normal training clips nor injected domain context at inference time.
- Outperforms frontier VLMs, agentic frameworks, and fine-tuned traditional methods across three IVAD datasets.
Sources (1)
O-VAD: Industrial Video Anomaly Detection through Object-Centric Tracking and Reasoning
TL;DR - O-VAD is a training-free, object-centric framework that tracks objects through industrial videos and reasons over their changing states to detect anomalies. It aims to improve detection under complex physical and procedural constraints without retraining or domain-specific knowledge.
- Models spatial-temporal object dynamics and transformations as state trajectories.
- Identifies anomalous objects in grounded video frames and generates interpretable reports.
- Requires neither normal training clips nor injected domain context at inference time.
- Outperforms frontier VLMs, agentic frameworks, and fine-tuned traditional methods across three IVAD datasets.