🛰️ Daily AI Frontier
‹ back to 2026-07-16

SceneBind: Binding What and Where Across Vision, Audio and Language

Research Multimodal & Generative

Ranking

Overall 79
Content 95
Popularity 41

Observed public metrics from 1 member.

Merged summary

TL;DR - SceneBind is an omni-modal encoder that adds explicit 3D spatial structure to scene representations across vision, audio, and language, bridging the gap between knowing what is present and where it is. It matters because most omni-modal models capture semantics but lack spatial grounding, which is essential for realistic scene understanding.

  • Represents each scene as a semantic-spatial entity: a global semantic embedding plus object-centric slots encoding object semantics, spatial attributes, and uncertainty.
  • Introduces "SceneBind Matching," combining global scene similarity with object-level alignment to support cross-modal scene retrieval and object grounding.
  • Curates a new real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for aligning semantic and spatial signals across modalities.
  • Works atop large pretrained semantic encoders with only lightweight added spatial tokens, claiming state-of-the-art scene/spatial retrieval and strong zero-shot transfer to audio-visual localization.

Sources (1)

SceneBind: Binding What and Where Across Vision, Audio and Language

arXiv cs.CV Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman 2026-07-16 arXiv:2607.15265
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-08-12 14:36:08.071558 UTC

TL;DR - SceneBind is an omni-modal encoder that adds explicit 3D spatial structure to scene representations across vision, audio, and language, bridging the gap between knowing what is present and where it is. It matters because most omni-modal models capture semantics but lack spatial grounding, which is essential for realistic scene understanding.

  • Represents each scene as a semantic-spatial entity: a global semantic embedding plus object-centric slots encoding object semantics, spatial attributes, and uncertainty.
  • Introduces "SceneBind Matching," combining global scene similarity with object-level alignment to support cross-modal scene retrieval and object grounding.
  • Curates a new real-world binaural audio-visual dataset with structured semantic and spatial annotations, plus a training protocol for aligning semantic and spatial signals across modalities.
  • Works atop large pretrained semantic encoders with only lightweight added spatial tokens, claiming state-of-the-art scene/spatial retrieval and strong zero-shot transfer to audio-visual localization.
item →