Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
Merged summary
TL;DR - A unified framework generates full-length vocal, instrumental, and cover songs by combining hierarchical autoregressive planning with flow-matching audio rendering. It targets both long-range musical structure and high-fidelity output.
- A semantic-aware tokenizer represents audio using eight-codebook RVQ tokens.
- A hierarchical autoregressive model plans complete songs, while FullDiT renders them through flow matching in a continuous VAE latent space.
- A two-level melody module preserves reference melodies during cover-song generation.
- Reward-based post-training includes DPO, GRPO, OPD, and flow-based GRPO; evaluations report competitive multilingual performance.
Sources (1)
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
TL;DR - A unified framework generates full-length vocal, instrumental, and cover songs by combining hierarchical autoregressive planning with flow-matching audio rendering. It targets both long-range musical structure and high-fidelity output.
- A semantic-aware tokenizer represents audio using eight-codebook RVQ tokens.
- A hierarchical autoregressive model plans complete songs, while FullDiT renders them through flow matching in a continuous VAE latent space.
- A two-level melody module preserves reference melodies during cover-song generation.
- Reward-based post-training includes DPO, GRPO, OPD, and flow-based GRPO; evaluations report competitive multilingual performance.