🛰️ Daily AI Frontier
‹ back to 2026-07-22

Context-structured Video Anomaly Detection with Large Vision-Language Models

arXiv cs.CV Multimodal & Generative Dongjun Kim, Changjae Oh, Andrea Cavallaro, Jeonghoon Mo 2026-07-21

TL;DR - CSI-VAD is a training-free video anomaly detector that uses large vision-language models to analyze environment, object, and temporal contexts separately. This structured approach improves detection over holistic inference without predefined anomaly prompts or dataset-specific tuning.

  • Decomposes videos into three context-specific branches: environment, objects, and time.
  • Grounds anomaly judgments solely in visual cues from each context.
  • Avoids costly anomaly annotations, handcrafted text prompts, and dataset-specific tuning.
  • Outperforms the direct holistic baseline and is competitive with existing methods on UCF-Crime and UBnormal.

view merged work →