WanSong v1.0 Technical Report
- Published
- Source
- arXiv
- Paper number
- 647
- Field
- Computer Vision
- arXiv ID
- 2607.14749
Key points
- It abandons the autoregressive pipeline and uses pure diffusion to generate a five-minute song with vocals and background music as dual stems in a single pass.
- Dual-stem tokens model vocals and accompaniment separately, so both can be generated at high quality without the usual quality drop from classifier-free guidance.
- A 1024-compression VAE, which is half the compression of the previous 2048 setting, preserves fine acoustic detail and improves SI-SDR by 2.86 dB.
- RLHF is used to optimize musicality, lyric accuracy, and prompt alignment, yielding a musicality score of 5.49 compared with 4.18 for Suno and 3.83 for Mureka.
- Stable long-song generation is achieved through 6 million hours of multilingual data and three-stage curriculum training from 90 seconds to 300 seconds to SFT.
Paper links
External research summaries. These are not HDATF publications or measured product results.