RT by @huggingface: MiniMax H3 just dropped on Hugging Face text-to-video, image-to-video…
TL;DR - MiniMax released H3, a 33B-parameter video generation model, on Hugging Face with text-to-video, image-to-video, and reference-to-video modes that all produce synchronized audio. It matters because open-weight video+audio generation at consumer-GPU scale narrows the gap with closed commercial video models.
- Three conditioning modes in one model: text-to-video, image-to-video, and reference-to-video, each generating video with accompanying audio rather than silent clips.
- 33B parameters, but the announcement claims it is runnable on consumer GPUs — implying quantization/offloading support rather than datacenter-only inference.
- Ships with day-one ecosystem integration: 🧨 diffusers and ComfyUI support, plus downloadable weights and a hosted demo app on Hugging Face Spaces.
- Note: this is a distribution/launch post from Hugging Face; no benchmarks, training details, or license terms are given in the content provided.