🛰️ Daily AI Frontier
‹ back to 2026-08-29

Omni-Interactive Universal Embedder

Research Multimodal & Generative

Ranking

Overall 79
Content 95
Popularity 42

Observed public metrics from 1 member.

Representative image for Omni-Interactive Universal Embedder

Merged summary

TL;DR - OmniUE is a universal embedder that maps text, video, and audio into a unified space while supporting queries conditioned on text, visual regions, and audio spans. It substantially improves several interactive retrieval benchmarks, suggesting a path toward more flexible any-to-any multimodal search.

  • Uses learnable tokens and intermediate omni-LLM representations to generate user-conditioned embeddings across modalities.
  • Adds visual and audio segmenters to incorporate selected regions of interest and temporal audio spans into retrieval queries.
  • Introduces OmniCHOIR, a benchmark for compositional audio retrieval using multimodal inputs and interaction prompts.
  • Reports average gains over state-of-the-art baselines of 10.5% on MMEB-v2-video, 1.1% on MAEB, 83.7% on SCaR, and 24.1% on OmniCHOIR.

Sources (1)

Omni-Interactive Universal Embedder

arXiv cs.AI Wei-Yao Wang, Kazuya Tateishi, Shuyang Cui, Christian Simon, Takashi Shibuya, Shusuke Takahashi, Yuki Mitsufuji 2026-08-27 arXiv:2608.27044
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-25 14:26:24.127021 UTC

TL;DR - OmniUE is a universal embedder that maps text, video, and audio into a unified space while supporting queries conditioned on text, visual regions, and audio spans. It substantially improves several interactive retrieval benchmarks, suggesting a path toward more flexible any-to-any multimodal search.

  • Uses learnable tokens and intermediate omni-LLM representations to generate user-conditioned embeddings across modalities.
  • Adds visual and audio segmenters to incorporate selected regions of interest and temporal audio spans into retrieval queries.
  • Introduces OmniCHOIR, a benchmark for compositional audio retrieval using multimodal inputs and interaction prompts.
  • Reports average gains over state-of-the-art baselines of 10.5% on MMEB-v2-video, 1.1% on MAEB, 83.7% on SCaR, and 24.1% on OmniCHOIR.
item →