🛰️ Daily AI Frontier
‹ back to 2026-07-16

SceneBind: Binding What and Where Across Vision, Audio and Language

arXiv cs.CV Multimodal & Generative Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman 2026-07-16

TL;DR - SceneBind is an omni-modal encoder that adds explicit 3D spatial structure to scene representations across vision, audio, and language, bridging the gap between knowing what is present and where it is. It matters because most omni-modal models capture semantics but lack spatial grounding, which is essential for realistic scene understanding.

  • Represents each scene as a semantic-spatial entity: a global semantic embedding plus object-centric slots encoding object semantics, spatial attributes, and uncertainty.
  • Introduces "SceneBind Matching," combining global scene similarity with object-level alignment to support cross-modal scene retrieval and object grounding.
  • Curates a new real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for aligning semantic and spatial signals across modalities.
  • Works atop large pretrained semantic encoders with only lightweight added spatial tokens, claiming state-of-the-art scene/spatial retrieval and strong zero-shot transfer to audio-visual localization.

view merged work →