🛰️ Daily AI Frontier
‹ back to 2026-07-23

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

Research Multimodal & Generative

Ranking

Overall 69
Content 80
Popularity 43

Observed public metrics from 1 member.

Merged summary

TL;DR - PercepCap is a video-captioning framework that explicitly generates object trajectories and temporal events before writing a caption. This makes perceptual errors more interpretable while improving caption quality over the Qwen3-VL baseline.

  • Uses a “perceive-then-describe” chain that conditions captions on explicit spatiotemporal evidence.
  • Combines supervised fine-tuning with perception-grounded reinforcement learning.
  • Constructs caption-aligned training data by grounding mentioned objects and events with bounding boxes and timestamps.
  • Reports consistent gains in direct-captioning and caption-to-QA evaluations.

Sources (1)

PercepCap: Video Captioner with Structured Spatio-Temporal Perception

arXiv cs.CV Yifan Xu, Zihao Wang, Zhixiao Wang, Jiaming Zhang, Yichun Yang, Desen Meng, Yuanxing Zhang, Pengfei Wan, Limin Wang 2026-07-22 arXiv:2607.20389
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-15 14:30:42.031198 UTC

TL;DR - PercepCap is a video-captioning framework that explicitly generates object trajectories and temporal events before writing a caption. This makes perceptual errors more interpretable while improving caption quality over the Qwen3-VL baseline.

  • Uses a “perceive-then-describe” chain that conditions captions on explicit spatiotemporal evidence.
  • Combines supervised fine-tuning with perception-grounded reinforcement learning.
  • Constructs caption-aligned training data by grounding mentioned objects and events with bounding boxes and timestamps.
  • Reports consistent gains in direct-captioning and caption-to-QA evaluations.
item →