🛰️ Daily AI Frontier
‹ back to 2026-08-10

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

Research Multimodal & Generative

Ranking

Overall 66
Content 75
Popularity 43

Observed public metrics from 1 member.

Representative image for I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

Merged summary

TL;DR - An arXiv cs.CV paper introduces Identity-conditioned Queries (ICQ), a video reasoning task where a model must jointly interpret a video and a reference image of a person, plus the ISYV suite (benchmark, training data, model) to support it. It matters because current video-language benchmarks assume simple video-text inputs and don't test identity grounding or person-centric temporal reasoning.

  • ISYV-Bench: 1,377 real-world complex videos with 1,377 QA pairs, organized into six difficulty levels ranging from identity recognition to causal reasoning.
  • ISYV-75K: 75K training samples built via automated annotation, multi-stage verification, and manual review.
  • ISYV-Framework: an ICQ-oriented model and training strategy that learns to exploit informative video shots without shot-level annotations.
  • Reported findings: both closed- and open-source MLLMs struggle on the benchmark, notably on cross-domain identity matching and long-horizon tracking; ISYV-Model beats strong baselines and approaches closed-source performance in some aspects.

Sources (1)

I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning

arXiv cs.CV Shibo Gao, Chongxiao Wang, Chenglong Huang, Jie Ma, Haolin Shi, Fei Ding, Jing Li, Qiang Lyu, Yangyang Liu, Yang Liu, Jun Liu, Linlin Huang, Peipei Yang 2026-08-07 arXiv:2608.07417
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-02 14:24:33.397985 UTC

TL;DR - An arXiv cs.CV paper introduces Identity-conditioned Queries (ICQ), a video reasoning task where a model must jointly interpret a video and a reference image of a person, plus the ISYV suite (benchmark, training data, model) to support it. It matters because current video-language benchmarks assume simple video-text inputs and don't test identity grounding or person-centric temporal reasoning.

  • ISYV-Bench: 1,377 real-world complex videos with 1,377 QA pairs, organized into six difficulty levels ranging from identity recognition to causal reasoning.
  • ISYV-75K: 75K training samples built via automated annotation, multi-stage verification, and manual review.
  • ISYV-Framework: an ICQ-oriented model and training strategy that learns to exploit informative video shots without shot-level annotations.
  • Reported findings: both closed- and open-source MLLMs struggle on the benchmark, notably on cross-domain identity matching and long-horizon tracking; ISYV-Model beats strong baselines and approaches closed-source performance in some aspects.
item →