DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
Ranking
Overall
87
Content
95
Popularity
69
Observed public metrics from 1 member.
Merged summary
TL;DR - DreamX-Creator 1.0 is a compact 7B model that jointly generates synchronized audio and video from a first frame and text prompt. Its released generator and one-step 2K refiner aim to make unified, high-resolution audio-video research more accessible.
- Jointly denoises modality-specific audio and video streams, coupling them later via Gated Cross-Modal Attention.
- Uses a unified data pipeline to filter temporally coherent clips, generate multimodal annotations, and organize capability-focused training pools.
- Combines progressive pretraining and high-quality fine-tuning with reinforcement learning that routes modality-aware feedback to audio, video, and cross-modal streams.
- Produces high-resolution output through an autoregressive 2K refinement pipeline distilled to one denoising evaluation per temporal chunk.
Sources (1)
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
Public signals
Hugging Face upvotes 95
TL;DR - DreamX-Creator 1.0 is a compact 7B model that jointly generates synchronized audio and video from a first frame and text prompt. Its released generator and one-step 2K refiner aim to make unified, high-resolution audio-video research more accessible.
- Jointly denoises modality-specific audio and video streams, coupling them later via Gated Cross-Modal Attention.
- Uses a unified data pipeline to filter temporally coherent clips, generate multimodal annotations, and organize capability-focused training pools.
- Combines progressive pretraining and high-quality fine-tuning with reinforcement learning that routes modality-aware feedback to audio, video, and cross-modal streams.
- Produces high-resolution output through an autoregressive 2K refinement pipeline distilled to one denoising evaluation per temporal chunk.