Qwen-Audio-3.0-Gen-Preview Technical Report

Published
Source
arXiv
Paper number
760
Field
Research
arXiv ID
2607.27011

Key points

  • It unifies speech, music, sound effects, and mixtures into a single generation path with a non-autoregressive structure.
  • It designs a shared VAE that compresses 48 kHz stereo into a 25 Hz latent representation.
  • It uses only about 10 percent of the music data used by specialized prior models, yet still leads in 3 of the 7 SongBench tasks.
  • On Seed-TTS-Eval, it achieves the highest speaker similarity, and it also outperforms Seed-Audio-1.0 in multi-speaker consistency.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)