🛰️ Daily AI Frontier
‹ back to 2026-07-23

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Research Multimodal & Generative

Merged summary

TL;DR - PercepCap is a video-captioning framework that explicitly generates object trajectories and temporal events before writing a caption. This makes perceptual errors more interpretable while improving caption quality over the Qwen3-VL baseline.

  • Uses a “perceive-then-describe” chain that conditions captions on explicit spatiotemporal evidence.
  • Combines supervised fine-tuning with perception-grounded reinforcement learning.
  • Constructs caption-aligned training data by grounding mentioned objects and events with bounding boxes and timestamps.
  • Reports consistent gains in direct-captioning and caption-to-QA evaluations.

Sources (1)

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

arXiv cs.CV Yifan Xu, Zihao Wang, Zhixiao Wang, Jiaming Zhang, Yichun Yang, Desen Meng, Yuanxing Zhang, Pengfei Wan, Limin Wang 2026-07-22 arXiv:2607.20389

TL;DR - PercepCap is a video-captioning framework that explicitly generates object trajectories and temporal events before writing a caption. This makes perceptual errors more interpretable while improving caption quality over the Qwen3-VL baseline.

  • Uses a “perceive-then-describe” chain that conditions captions on explicit spatiotemporal evidence.
  • Combines supervised fine-tuning with perception-grounded reinforcement learning.
  • Constructs caption-aligned training data by grounding mentioned objects and events with bounding boxes and timestamps.
  • Reports consistent gains in direct-captioning and caption-to-QA evaluations.
item →