SceneBind: Binding What and Where Across Vision, Audio and Language
Merged summary
TL;DR - SceneBind is an omni-modal encoder that adds explicit 3D spatial structure to scene representations across vision, audio, and language, bridging the gap between knowing what is present and where it is. It matters because most omni-modal models capture semantics but lack spatial grounding, which is essential for realistic scene understanding.
- Represents each scene as a semantic-spatial entity: a global semantic embedding plus object-centric slots encoding object semantics, spatial attributes, and uncertainty.
- Introduces "SceneBind Matching," combining global scene similarity with object-level alignment to support cross-modal scene retrieval and object grounding.
- Curates a new real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for aligning semantic and spatial signals across modalities.
- Works atop large pretrained semantic encoders with only lightweight added spatial tokens, claiming state-of-the-art scene/spatial retrieval and strong zero-shot transfer to audio-visual localization.
Sources (1)
SceneBind: Binding What and Where Across Vision, Audio and Language
TL;DR - SceneBind is an omni-modal encoder that adds explicit 3D spatial structure to scene representations across vision, audio, and language, bridging the gap between knowing what is present and where it is. It matters because most omni-modal models capture semantics but lack spatial grounding, which is essential for realistic scene understanding.
- Represents each scene as a semantic-spatial entity: a global semantic embedding plus object-centric slots encoding object semantics, spatial attributes, and uncertainty.
- Introduces "SceneBind Matching," combining global scene similarity with object-level alignment to support cross-modal scene retrieval and object grounding.
- Curates a new real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for aligning semantic and spatial signals across modalities.
- Works atop large pretrained semantic encoders with only lightweight added spatial tokens, claiming state-of-the-art scene/spatial retrieval and strong zero-shot transfer to audio-visual localization.