🛰️ Daily AI Frontier
‹ back to 2026-08-02

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

arXiv cs.CV Multimodal & Generative Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang 2026-07-30
Representative image for RefCaptioner: Multi-Reference Image-Grounded Video Captioning

TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.

  • Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
  • Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
  • Trains on 20,000 videos paired with 171,354 reference images.
  • Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.

view merged work →