RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection
Ranking
Overall
59
Content
65
Popularity
45
Observed public metrics from 1 member.
Merged summary
TL;DR - RDVSv2 is a large-scale RGB-D video salient-object-detection benchmark with dense, eye-tracking-guided annotations. It also introduces a parameter-efficient SAM2 baseline that achieves state-of-the-art results across RDVSv2 and existing benchmarks.
- Contains 249 stereoscopic video sequences and 29,077 annotated frames.
- Provides stereo-derived depth maps and frame-level salient-object masks.
- Covers more diverse and challenging scenarios than prior RGB-D VSOD datasets.
- The baseline jointly encodes RGB, depth, and optical-flow cues by fine-tuning the SAM2 encoder with PEFT.
Sources (1)
RDVSv2: A Large-scale Benchmark for RGB-D Video Salient Object Detection
Public signals
Semantic Scholar citations 0 · Semantic Scholar influential citations 0
TL;DR - RDVSv2 is a large-scale RGB-D video salient-object-detection benchmark with dense, eye-tracking-guided annotations. It also introduces a parameter-efficient SAM2 baseline that achieves state-of-the-art results across RDVSv2 and existing benchmarks.
- Contains 249 stereoscopic video sequences and 29,077 annotated frames.
- Provides stereo-derived depth maps and frame-level salient-object masks.
- Covers more diverse and challenging scenarios than prior RGB-D VSOD datasets.
- The baseline jointly encodes RGB, depth, and optical-flow cues by fine-tuning the SAM2 encoder with PEFT.