Qwen-Music Technical Report

Published
Source
arXiv
Paper number
606
Field
Research
arXiv ID
2607.11699

Key points

  • It compresses an entire song into a 25 Hz single-codebook Music Semantic Token so the LLM has room to predict the sequence.
  • The Melody-CoT mechanism plans melody as an explicit intermediate representation before token generation, which improves musicality and structural consistency.
  • Qwen-Music-Render combines a semantic-conditioned DiT, Spec-VAE, and Band-Mode Refiner to synthesize 48 kHz stereo audio.
  • It is trained on more than 5 million hours of multilingual music data with a curriculum by quality level and multi-stage post-training from SFT to DPO to GSPO.
  • In blind tests by expert evaluators, it achieves a 66.7% win rate against MiniMax 2.6 and a 58.3% win rate against Mureka V8.
  • It reaches state of the art on 13 of 16 objective metrics, including SongBench, SongEval, and AudioBox-Aesthetic.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)