🛰️ Daily AI Frontier
‹ back to 2026-08-29

Omni-Interactive Universal Embedder

arXiv cs.AI Multimodal & Generative Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji 2026-08-27
Representative image for Omni-Interactive Universal Embedder

TL;DR - OmniUE is a universal embedder that maps text, video, and audio into a unified space while supporting queries conditioned on text, visual regions, and audio spans. It substantially improves several interactive retrieval benchmarks, suggesting a path toward more flexible any-to-any multimodal search.

  • Uses learnable tokens and intermediate omni-LLM representations to generate user-conditioned embeddings across modalities.
  • Adds visual and audio segmenters to incorporate selected regions of interest and temporal audio spans into retrieval queries.
  • Introduces OmniCHOIR, a benchmark for compositional audio retrieval using multimodal inputs and interaction prompts.
  • Reports average gains over state-of-the-art baselines of 10.5% on MMEB-v2-video, 1.1% on MAEB, 83.7% on SCaR, and 24.1% on OmniCHOIR.

view merged work →