Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
Merged summary
TL;DR - A frozen large view-synthesis model can propagate panoptic labels across camera views, enabling 3D scene segmentation without explicit reconstruction or segmentation-specific model training. The approach matches reconstruction-based segmentation on ScanNet while improving novel-view synthesis and transferring to Replica without fine-tuning.
- Encodes input-view panoptic labels as binary channels and renders target-view segmentations through learned cross-view attention.
- Shows that spatial correspondences learned solely from RGB supervision generalize to view-independent per-pixel labels.
- Matches Gaussian-based segmentation methods on ScanNet while exceeding their novel-view synthesis quality by more than 7 dB.
- Outperforms those approaches on Replica without dataset-specific fine-tuning.
Sources (1)
Extending a Large View Synthesis Model for Multi-view Panoptic Segmentation
TL;DR - A frozen large view-synthesis model can propagate panoptic labels across camera views, enabling 3D scene segmentation without explicit reconstruction or segmentation-specific model training. The approach matches reconstruction-based segmentation on ScanNet while improving novel-view synthesis and transferring to Replica without fine-tuning.
- Encodes input-view panoptic labels as binary channels and renders target-view segmentations through learned cross-view attention.
- Shows that spatial correspondences learned solely from RGB supervision generalize to view-independent per-pixel labels.
- Matches Gaussian-based segmentation methods on ScanNet while exceeding their novel-view synthesis quality by more than 7 dB.
- Outperforms those approaches on Replica without dataset-specific fine-tuning.