Qwen-Audio-VAE Technical Report

Published
Source
arXiv
Paper number
612
Field
Research
arXiv ID
2607.11738

Key points

  • We reduce downstream DiT training cost by using a 12.5 Hz low-bit-rate latent space.
  • An asymmetric encoder-decoder with latency-aware encoder pruning encodes 32 minutes in 541 ms, a 3.62x speedup.
  • Trained on 5 million hours of multi-domain audio, it generalizes across speech, music, and sound-effect domains.
  • Four discriminators, multi-period, multi-resolution STFT, multi-scale STFT, and sub-band CQT, provide adversarial supervision.
  • It achieves robust reconstruction quality across domains on the public LibriSpeech, AudioCaps, and SongDescriber benchmarks.
  • It outperforms the previous best VAE on STFT distance, SI-SDR, and related metrics.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)