🛰️ Daily AI Frontier
‹ back to 2026-08-02

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Research Multimodal & Generative

Ranking

Overall 80
Content 85
Popularity 69

Observed public metrics from 1 member.

Representative image for RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Merged summary

TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.

  • Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
  • Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
  • Trains on 20,000 videos paired with 171,354 reference images.
  • Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.

Sources (1)

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

arXiv cs.CV Tengfei Liu, Yang Shi, Yuran Wang, Xiaohan Zhang, Yuqing Wen, Yuqi Tang, Qixun Wang, Zhuoran Zhang, Xuanyu Zhu, Weihong Lin, Xinlei Yu, Yujie Wei, Xinwei Long, Fengxiang Wang, Xinlong Chen, Yue Ding, Jialu Chen, Haotian Wang, Yuanxing Zhang 2026-07-30 arXiv:2607.28509
Public signals Hugging Face upvotes 30 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · Upvotes 30 OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-31 14:29:56.538590 UTC

TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.

  • Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
  • Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
  • Trains on 20,000 videos paired with 171,354 reference images.
  • Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.
item →