ECCV 2026|Pistachio 全面开源:用可控视频生成,重新构建视频异常检测与理解 Benchmark
TL;DR - ECCV 2026 paper Pistachio introduces an open synthetic benchmark for video anomaly detection and understanding, built with controllable long-form video generation. It expands anomaly diversity and exposes substantial weaknesses in current VAD and multimodal models.
- Pistachio-VAD contains 4,962 videos across 31 anomaly classes; Pistachio-VAU adds 1,385 videos with event- and video-level descriptions.
- A hierarchical pipeline generates temporally connected story segments, then derives semantic annotations from the storyline and corrects them against generated footage using a VLM.
- The best reported VAD results reached 83.7% AUC and 71.9% AP, while the best VAU model averaged only 29.39% F1.
- The dataset, generation pipeline, features, and adaptations for nine VAD baselines are open source; combining Pistachio with 10% of UCF-Crime outperformed training on the complete real dataset alone.