Qwen-Audio-VAE Technical Report
- Published
- Source
- arXiv
- Paper number
- 612
- Field
- Research
- arXiv ID
- 2607.11738
Key points
- We reduce downstream DiT training cost by using a 12.5 Hz low-bit-rate latent space.
- An asymmetric encoder-decoder with latency-aware encoder pruning encodes 32 minutes in 541 ms, a 3.62x speedup.
- Trained on 5 million hours of multi-domain audio, it generalizes across speech, music, and sound-effect domains.
- Four discriminators, multi-period, multi-resolution STFT, multi-scale STFT, and sub-band CQT, provide adversarial supervision.
- It achieves robust reconstruction quality across domains on the public LibriSpeech, AudioCaps, and SongDescriber benchmarks.
- It outperforms the previous best VAE on STFT distance, SI-SDR, and related metrics.
Paper links
External research summaries. These are not HDATF publications or measured product results.