Qwen-Music Technical Report
- Published
- Source
- arXiv
- Paper number
- 606
- Field
- Research
- arXiv ID
- 2607.11699
Key points
- It compresses an entire song into a 25 Hz single-codebook Music Semantic Token so the LLM has room to predict the sequence.
- The Melody-CoT mechanism plans melody as an explicit intermediate representation before token generation, which improves musicality and structural consistency.
- Qwen-Music-Render combines a semantic-conditioned DiT, Spec-VAE, and Band-Mode Refiner to synthesize 48 kHz stereo audio.
- It is trained on more than 5 million hours of multilingual music data with a curriculum by quality level and multi-stage post-training from SFT to DPO to GSPO.
- In blind tests by expert evaluators, it achieves a 66.7% win rate against MiniMax 2.6 and a 58.3% win rate against Mureka V8.
- It reaches state of the art on 13 of 16 objective metrics, including SongBench, SongEval, and AudioBox-Aesthetic.
Paper links
External research summaries. These are not HDATF publications or measured product results.