RefCaptioner: Multi-Reference Image-Grounded Video Captioning
TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.
- Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
- Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
- Trains on 20,000 videos paired with 171,354 reference images.
- Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.