Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering

Published
Source
arXiv
Paper number
695
Field
AI / General
arXiv ID
2607.20253

Key points

  • It supports three tasks in one framework: lyric-to-song generation, instrumental generation, and cover-song generation.
  • It reaches an Elo score of 1,129 on the Artificial Analysis Music with Vocals leaderboard, which ties for second to third place and is only 12 points behind first.
  • In tests on 500 songs across 8 genres and 5 languages, it gets the top score on 15 of 18 metrics.
  • FullDiT performs flow matching on the entire song rather than on chunks, which preserves global consistency.
  • Post-training with DPO, GRPO, and OPD improves both musicality and rendering quality.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)