全球AI视频榜单第一梯队再添中国力量:智象发布首款物理规律导向视频模型
TL;DR - HiDream.ai launched HD-V1, a natively multimodal video model that generates 5–20-second, 1080p videos from text, images, or video. It emphasizes physical realism, autonomous narrative planning, and synchronized audio-video generation, ranking fourth on Artificial Analysis’s image-to-video-with-audio leaderboard and eighth on Arena.ai’s image-to-video benchmark.
- A multimodal intent module converts natural-language prompts into structured plans covering shots, characters, actions, camera work, dialogue, and sound.
- The model jointly handles text, video, and audio, targeting consistent motion, synchronized sound effects and speech, and coherent multi-shot narratives.
- Its pipeline combines upfront planning, joint generation, and post-training with diffusion reinforcement learning and a multimodal reward model.
- HD-V1 can select a 5–20-second duration based on the action and narrative rather than mechanically filling a fixed user-specified runtime.