🛰️ Daily AI Frontier
‹ back to 2026-08-10

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

arXiv cs.CV Multimodal & Generative Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang 2026-08-07
Representative image for I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

TL;DR - An arXiv cs.CV paper introduces Identity-conditioned Queries (ICQ), a video reasoning task where a model must jointly interpret a video and a reference image of a person, plus the ISYV suite (benchmark, training data, model) to support it. It matters because current video-language benchmarks assume simple video-text inputs and don't test identity grounding or person-centric temporal reasoning.

  • ISYV-Bench: 1,377 real-world complex videos with 1,377 QA pairs, organized into six difficulty levels ranging from identity recognition to causal reasoning.
  • ISYV-75K: 75K training samples built via automated annotation, multi-stage verification, and manual review.
  • ISYV-Framework: an ICQ-oriented model and training strategy that learns to exploit informative video shots without shot-level annotations.
  • Reported findings: both closed- and open-source MLLMs struggle on the benchmark, notably on cross-domain identity matching and long-horizon tracking; ISYV-Model beats strong baselines and approaches closed-source performance in some aspects.

view merged work →