🛰️ Daily AI Frontier
‹ back to 2026-08-20

ICML 2026 | 从专才到通才,OmniShow极简干预统一多模态视频生成

Research Multimodal & Generative

Ranking

Overall 80
Content 85
Popularity 68

Observed public metrics from 1 member.

Representative image for ICML 2026 | 从专才到通才,OmniShow极简干预统一多模态视频生成

Merged summary

TL;DR - OmniShow is an ICML 2026 paper proposing a unified 12.3B video model conditioned jointly on text, reference images, audio, and pose. It preserves the base model’s generation priors through minimal architectural changes while achieving strong multimodal control and competitive benchmark results.

  • Visual conditions reuse Waver 1.0’s native channel-concatenation path, with pseudo-frame reference tokens and a reconstruction loss improving identity and object fidelity.
  • Gated local-context attention aligns audio with video frames; near-zero gate initialization limits disruption, and the added audio components increase parameters by only about 2.5%.
  • A decoupled-then-joint training strategy merges reference-video and audio-video specialists before joint refinement, producing zero-shot reference-and-audio control immediately after weight interpolation.
  • On HOIVG-Bench, OmniShow reports leading or highly competitive reference consistency, audiovisual synchronization, pose accuracy, and video-quality metrics across R2V, RA2V, and RP2V settings.

Sources (1)

ICML 2026 | 从专才到通才,OmniShow极简干预统一多模态视频生成

WeChat: PaperWeekly 2026-08-20 arXiv:2604.11804
Public signals Hugging Face upvotes 72
Providers: Hugging Face · Upvotes 72 OpenAlex · N/A Publisher · N/A Semantic Scholar · N/A X · N/A Fetched 2026-09-19 14:26:06.571727 UTC

TL;DR - OmniShow is an ICML 2026 paper proposing a unified 12.3B video model conditioned jointly on text, reference images, audio, and pose. It preserves the base model’s generation priors through minimal architectural changes while achieving strong multimodal control and competitive benchmark results.

  • Visual conditions reuse Waver 1.0’s native channel-concatenation path, with pseudo-frame reference tokens and a reconstruction loss improving identity and object fidelity.
  • Gated local-context attention aligns audio with video frames; near-zero gate initialization limits disruption, and the added audio components increase parameters by only about 2.5%.
  • A decoupled-then-joint training strategy merges reference-video and audio-video specialists before joint refinement, producing zero-shot reference-and-audio control immediately after weight interpolation.
  • On HOIVG-Bench, OmniShow reports leading or highly competitive reference consistency, audiovisual synchronization, pose accuracy, and video-quality metrics across R2V, RA2V, and RP2V settings.
item →