SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Published
Source
arXiv
Paper number
804
Field
Research
arXiv ID
2608.02023

Key points

  • Speech preprocessing, multilayer captions, quality refinement, and synthesis of rare situations produced approximately 70 million training records describing environments, speakers, and content.
  • It combined SwanVAE and a flow-based transformer with quality conditioning, Engram, and a unified mixture of experts to handle reference speech and natural-language instructions in one model.
  • In single-speaker evaluation with reference speech, it achieved timbre similarity of 0.95 and content error of 0.086, improving over the earlier SwanVoice.
  • In instruction evaluations, it ranked 1st in Chinese with APS 86.1 and tied for 1st in English with APS 84.2, and could generate speech, ambient sound, and sound effects in one waveform.
  • Weaknesses remain in emotional transitions in complex background music, multi-speaker scenes longer than 2 minutes, continuous control of a particular speaker's emotion, stress, and pauses, and role acting.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)