Context-structured Video Anomaly Detection with Large Vision-Language Models
Merged summary
TL;DR - CSI-VAD is a training-free video anomaly detector that uses large vision-language models to analyze environment, object, and temporal contexts separately. This structured approach improves detection over holistic inference without predefined anomaly prompts or dataset-specific tuning.
- Decomposes videos into three context-specific branches: environment, objects, and time.
- Grounds anomaly judgments solely in visual cues from each context.
- Avoids costly anomaly annotations, handcrafted text prompts, and dataset-specific tuning.
- Outperforms the direct holistic baseline and is competitive with existing methods on UCF-Crime and UBnormal.
Sources (1)
Context-structured Video Anomaly Detection with Large Vision-Language Models
TL;DR - CSI-VAD is a training-free video anomaly detector that uses large vision-language models to analyze environment, object, and temporal contexts separately. This structured approach improves detection over holistic inference without predefined anomaly prompts or dataset-specific tuning.
- Decomposes videos into three context-specific branches: environment, objects, and time.
- Grounds anomaly judgments solely in visual cues from each context.
- Avoids costly anomaly annotations, handcrafted text prompts, and dataset-specific tuning.
- Outperforms the direct holistic baseline and is competitive with existing methods on UCF-Crime and UBnormal.