🛰️ Daily AI Frontier
‹ back to 2026-07-22

Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval

arXiv cs.CV Multimodal & Generative Jihyun Lee, Cheol-Ho Cho, Woojin Jun, Woojin Jeong, Jae-Pil Heo 2026-07-21

TL;DR - Self-SiMS improves zero-shot video moment retrieval by generating and scoring temporal spans from within-video self-similarity rather than unreliable text-video similarity. It also uses query-aware multimodal LLM reasoning to refine alignment, achieving state-of-the-art benchmark performance.

  • Avoids dependence on large-scale text-to-temporal-span annotations.
  • Uses intrinsic video relationships to mitigate modality and language-style gaps.
  • Produces more robust span proposals and retrieval scores.
  • Adds query-aware MLLM reasoning to sharpen text-video alignment.

view merged work →