RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Ranking
Overall
80
Content
85
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.
- Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
- Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
- Trains on 20,000 videos paired with 171,354 reference images.
- Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.
Sources (1)
RefCaptioner: Multi-Reference Image-Grounded Video Captioning
Public signals
Hugging Face upvotes 30 · Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - RefCaptioner tackles multi-reference image-grounded video captioning, producing factual captions that bind phrases to specific reference images. It improves grounding while preserving general video-captioning performance.
- Uses mixed-data supervised fine-tuning and Hierarchical Coverage-Discounted GRPO.
- Targets reference selection, phrase-level binding, distractor rejection, and cross-reference consistency.
- Trains on 20,000 videos paired with 171,354 reference images.
- Introduces MRVBench for evaluating caption factuality and grounding on real and AI-generated videos.