🛰️ Daily AI Frontier
‹ back to 2026-07-22

Context-structured Video Anomaly Detection with Large Vision-Language Models

Research Multimodal & Generative

Ranking

Overall 57
Content 65
Popularity 40

Observed public metrics from 1 member.

Merged summary

TL;DR - CSI-VAD is a training-free video anomaly detector that uses large vision-language models to analyze environment, object, and temporal contexts separately. This structured approach improves detection over holistic inference without predefined anomaly prompts or dataset-specific tuning.

  • Decomposes videos into three context-specific branches: environment, objects, and time.
  • Grounds anomaly judgments solely in visual cues from each context.
  • Avoids costly anomaly annotations, handcrafted text prompts, and dataset-specific tuning.
  • Outperforms the direct holistic baseline and is competitive with existing methods on UCF-Crime and UBnormal.

Sources (1)

Context-structured Video Anomaly Detection with Large Vision-Language Models

arXiv cs.CV Dongjun Kim, Changjae Oh, Andrea Cavallaro, Jeonghoon Mo 2026-07-21 arXiv:2607.19077
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-16 14:17:07.805125 UTC

TL;DR - CSI-VAD is a training-free video anomaly detector that uses large vision-language models to analyze environment, object, and temporal contexts separately. This structured approach improves detection over holistic inference without predefined anomaly prompts or dataset-specific tuning.

  • Decomposes videos into three context-specific branches: environment, objects, and time.
  • Grounds anomaly judgments solely in visual cues from each context.
  • Avoids costly anomaly annotations, handcrafted text prompts, and dataset-specific tuning.
  • Outperforms the direct holistic baseline and is competitive with existing methods on UCF-Crime and UBnormal.
item →