WanSong v1.0 Technical Report

Published
Source
arXiv
Paper number
647
Field
Computer Vision
arXiv ID
2607.14749

Key points

  • It abandons the autoregressive pipeline and uses pure diffusion to generate a five-minute song with vocals and background music as dual stems in a single pass.
  • Dual-stem tokens model vocals and accompaniment separately, so both can be generated at high quality without the usual quality drop from classifier-free guidance.
  • A 1024-compression VAE, which is half the compression of the previous 2048 setting, preserves fine acoustic detail and improves SI-SDR by 2.86 dB.
  • RLHF is used to optimize musicality, lyric accuracy, and prompt alignment, yielding a musicality score of 5.49 compared with 4.18 for Suno and 3.83 for Mureka.
  • Stable long-song generation is achieved through 6 million hours of multilingual data and three-stage curriculum training from 90 seconds to 300 seconds to SFT.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)