🛰️ Daily AI Frontier
‹ back to 2026-07-23

Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

arXiv cs.SD Multimodal & Generative Junyu Dai, Xinyue Fan, Weiqin Li, Xiangang Li, Yunjia Li, Bin Ma, Yukun Ma, Chongjia Ni, Yufei Shi, Haoxu Wang, Menglin Wu, Jianwei Yu, Huaicheng Zhang, Han Zhao, Shengkui Zhao, Haina Zhu 2026-07-22

TL;DR - A unified framework generates full-length vocal, instrumental, and cover songs by combining hierarchical autoregressive planning with flow-matching audio rendering. It targets both long-range musical structure and high-fidelity output.

  • A semantic-aware tokenizer represents audio using eight-codebook RVQ tokens.
  • A hierarchical autoregressive model plans complete songs, while FullDiT renders them through flow matching in a continuous VAE latent space.
  • A two-level melody module preserves reference melodies during cover-song generation.
  • Reward-based post-training includes DPO, GRPO, OPD, and flow-based GRPO; evaluations report competitive multilingual performance.

view merged work →