Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
- Published
- Source
- arXiv
- Paper number
- 695
- Field
- AI / General
- arXiv ID
- 2607.20253
Key points
- It supports three tasks in one framework: lyric-to-song generation, instrumental generation, and cover-song generation.
- It reaches an Elo score of 1,129 on the Artificial Analysis Music with Vocals leaderboard, which ties for second to third place and is only 12 points behind first.
- In tests on 500 songs across 8 genres and 5 languages, it gets the top score on 15 of 18 metrics.
- FullDiT performs flow matching on the entire song rather than on chunks, which preserves global consistency.
- Post-training with DPO, GRPO, and OPD improves both musicality and rendering quality.
Paper links
External research summaries. These are not HDATF publications or measured product results.