Qwen-Audio-3.0-Gen-Preview Technical Report
- Published
- Source
- arXiv
- Paper number
- 760
- Field
- Research
- arXiv ID
- 2607.27011
Key points
- It unifies speech, music, sound effects, and mixtures into a single generation path with a non-autoregressive structure.
- It designs a shared VAE that compresses 48 kHz stereo into a 25 Hz latent representation.
- It uses only about 10 percent of the music data used by specialized prior models, yet still leads in 3 of the 7 SongBench tasks.
- On Seed-TTS-Eval, it achieves the highest speaker similarity, and it also outperforms Seed-Audio-1.0 in multi-speaker consistency.
Paper links
External research summaries. These are not HDATF publications or measured product results.