🛰️ Daily AI Frontier
‹ back to 2026-07-29

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

Research Multimodal & Generative

Ranking

Overall 59
Content 65
Popularity 45

Observed public metrics from 1 member.

Representative image for RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

Merged summary

TL;DR - RDVSv2 is a large-scale RGB-D video salient-object-detection benchmark with dense, eye-tracking-guided annotations. It also introduces a parameter-efficient SAM2 baseline that achieves state-of-the-art results across RDVSv2 and existing benchmarks.

  • Contains 249 stereoscopic video sequences and 29,077 annotated frames.
  • Provides stereo-derived depth maps and frame-level salient-object masks.
  • Covers more diverse and challenging scenarios than prior RGB-D VSOD datasets.
  • The baseline jointly encodes RGB, depth, and optical-flow cues by fine-tuning the SAM2 encoder with PEFT.

Sources (1)

RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection

arXiv cs.CV Tianyu Li, Jiahao He, Keren Fu, Qijun Zhao 2026-07-28 arXiv:2607.25392
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-28 14:34:21.017380 UTC

TL;DR - RDVSv2 is a large-scale RGB-D video salient-object-detection benchmark with dense, eye-tracking-guided annotations. It also introduces a parameter-efficient SAM2 baseline that achieves state-of-the-art results across RDVSv2 and existing benchmarks.

  • Contains 249 stereoscopic video sequences and 29,077 annotated frames.
  • Provides stereo-derived depth maps and frame-level salient-object masks.
  • Covers more diverse and challenging scenarios than prior RGB-D VSOD datasets.
  • The baseline jointly encodes RGB, depth, and optical-flow cues by fine-tuning the SAM2 encoder with PEFT.
item →