ICML 2026 | 从专才到通才,OmniShow极简干预统一多模态视频生成
Ranking
Overall
80
Content
85
Popularity
68
Observed public metrics from 1 member.
Merged summary
TL;DR - OmniShow is an ICML 2026 paper proposing a unified 12.3B video model conditioned jointly on text, reference images, audio, and pose. It preserves the base model’s generation priors through minimal architectural changes while achieving strong multimodal control and competitive benchmark results.
- Visual conditions reuse Waver 1.0’s native channel-concatenation path, with pseudo-frame reference tokens and a reconstruction loss improving identity and object fidelity.
- Gated local-context attention aligns audio with video frames; near-zero gate initialization limits disruption, and the added audio components increase parameters by only about 2.5%.
- A decoupled-then-joint training strategy merges reference-video and audio-video specialists before joint refinement, producing zero-shot reference-and-audio control immediately after weight interpolation.
- On HOIVG-Bench, OmniShow reports leading or highly competitive reference consistency, audiovisual synchronization, pose accuracy, and video-quality metrics across R2V, RA2V, and RP2V settings.
Sources (1)
ICML 2026 | 从专才到通才,OmniShow极简干预统一多模态视频生成
Public signals
Hugging Face upvotes 72
TL;DR - OmniShow is an ICML 2026 paper proposing a unified 12.3B video model conditioned jointly on text, reference images, audio, and pose. It preserves the base model’s generation priors through minimal architectural changes while achieving strong multimodal control and competitive benchmark results.
- Visual conditions reuse Waver 1.0’s native channel-concatenation path, with pseudo-frame reference tokens and a reconstruction loss improving identity and object fidelity.
- Gated local-context attention aligns audio with video frames; near-zero gate initialization limits disruption, and the added audio components increase parameters by only about 2.5%.
- A decoupled-then-joint training strategy merges reference-video and audio-video specialists before joint refinement, producing zero-shot reference-and-audio control immediately after weight interpolation.
- On HOIVG-Bench, OmniShow reports leading or highly competitive reference consistency, audiovisual synchronization, pose accuracy, and video-quality metrics across R2V, RA2V, and RP2V settings.