Omni-Interactive Universal Embedder
TL;DR - OmniUE is a universal embedder that maps text, video, and audio into a unified space while supporting queries conditioned on text, visual regions, and audio spans. It substantially improves several interactive retrieval benchmarks, suggesting a path toward more flexible any-to-any multimodal search.
- Uses learnable tokens and intermediate omni-LLM representations to generate user-conditioned embeddings across modalities.
- Adds visual and audio segmenters to incorporate selected regions of interest and temporal audio spans into retrieval queries.
- Introduces OmniCHOIR, a benchmark for compositional audio retrieval using multimodal inputs and interaction prompts.
- Reports average gains over state-of-the-art baselines of 10.5% on MMEB-v2-video, 1.1% on MAEB, 83.7% on SCaR, and 24.1% on OmniCHOIR.