AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
Ranking
Observed public metrics from 1 member.
Merged summary
TL;DR - AVCap is an arXiv preprint tackling detailed audio-video joint captioning with a new 100K-caption dataset, a reinforcement-learning recipe using detail-aware rewards, and a benchmark/metric for atomic-level evaluation. It matters because coarse rewards and missing fine-grained benchmarks have been the bottleneck for multimodal video understanding and generation.
- AVCap-100K: 100K temporally aligned, detail-rich audio-video caption pairs, addressing the scarcity of high-quality public audiovisual joint-caption data.
- Da-GRPO (Detail-Aware GRPO): a GRPO variant with finer-grained reward signals replacing the coarse rewards used in prior RL-based captioning work.
- Reported results: state-of-the-art among open-source models, matching or surpassing proprietary models on several evaluations (no specific numbers given in the abstract).
- AVCap-Bench / AVCap-Score: a dedicated benchmark and metric scoring captions at the atomic detail level; code, models, and data released on Hugging Face.
Sources (1)
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
TL;DR - AVCap is an arXiv preprint tackling detailed audio-video joint captioning with a new 100K-caption dataset, a reinforcement-learning recipe using detail-aware rewards, and a benchmark/metric for atomic-level evaluation. It matters because coarse rewards and missing fine-grained benchmarks have been the bottleneck for multimodal video understanding and generation.
- AVCap-100K: 100K temporally aligned, detail-rich audio-video caption pairs, addressing the scarcity of high-quality public audiovisual joint-caption data.
- Da-GRPO (Detail-Aware GRPO): a GRPO variant with finer-grained reward signals replacing the coarse rewards used in prior RL-based captioning work.
- Reported results: state-of-the-art among open-source models, matching or surpassing proprietary models on several evaluations (no specific numbers given in the abstract).
- AVCap-Bench / AVCap-Score: a dedicated benchmark and metric scoring captions at the atomic detail level; code, models, and data released on Hugging Face.