I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - An arXiv cs.CV paper introduces Identity-conditioned Queries (ICQ), a video reasoning task where a model must jointly interpret a video and a reference image of a person, plus the ISYV suite (benchmark, training data, model) to support it. It matters because current video-language benchmarks assume simple video-text inputs and don't test identity grounding or person-centric temporal reasoning.
- ISYV-Bench: 1,377 real-world complex videos with 1,377 QA pairs, organized into six difficulty levels ranging from identity recognition to causal reasoning.
- ISYV-75K: 75K training samples built via automated annotation, multi-stage verification, and manual review.
- ISYV-Framework: an ICQ-oriented model and training strategy that learns to exploit informative video shots without shot-level annotations.
- Reported findings: both closed- and open-source MLLMs struggle on the benchmark, notably on cross-domain identity matching and long-horizon tracking; ISYV-Model beats strong baselines and approaches closed-source performance in some aspects.
Sources (1)
I Seek You in Videos: Identity-Conditioned Queries for Person-Centric Video Reasoning
TL;DR - An arXiv cs.CV paper introduces Identity-conditioned Queries (ICQ), a video reasoning task where a model must jointly interpret a video and a reference image of a person, plus the ISYV suite (benchmark, training data, model) to support it. It matters because current video-language benchmarks assume simple video-text inputs and don't test identity grounding or person-centric temporal reasoning.
- ISYV-Bench: 1,377 real-world complex videos with 1,377 QA pairs, organized into six difficulty levels ranging from identity recognition to causal reasoning.
- ISYV-75K: 75K training samples built via automated annotation, multi-stage verification, and manual review.
- ISYV-Framework: an ICQ-oriented model and training strategy that learns to exploit informative video shots without shot-level annotations.
- Reported findings: both closed- and open-source MLLMs struggle on the benchmark, notably on cross-domain identity matching and long-horizon tracking; ISYV-Model beats strong baselines and approaches closed-source performance in some aspects.