🛰️ Daily AI Frontier
‹ back to 2026-09-01

DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

arXiv cs.CV Multimodal & Generative Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang, Bingze Song, Jiachen Lei, Ruimin Lin, Jiahong Wu, Xiangxiang Chu 2026-08-31
Representative image for DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

TL;DR - DreamX-Creator 1.0 is a compact 7B model that jointly generates synchronized audio and video from a first frame and text prompt. Its released generator and one-step 2K refiner aim to make unified, high-resolution audio-video research more accessible.

  • Jointly denoises modality-specific audio and video streams, coupling them later via Gated Cross-Modal Attention.
  • Uses a unified data pipeline to filter temporally coherent clips, generate multimodal annotations, and organize capability-focused training pools.
  • Combines progressive pretraining and high-quality fine-tuning with reinforcement learning that routes modality-aware feedback to audio, video, and cross-modal streams.
  • Produces high-resolution output through an autoregressive 2K refinement pipeline distilled to one denoising evaluation per temporal chunk.

view merged work →