SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
- Published
- Source
- arXiv
- Paper number
- 804
- Field
- Research
- arXiv ID
- 2608.02023
Key points
- Speech preprocessing, multilayer captions, quality refinement, and synthesis of rare situations produced approximately 70 million training records describing environments, speakers, and content.
- It combined SwanVAE and a flow-based transformer with quality conditioning, Engram, and a unified mixture of experts to handle reference speech and natural-language instructions in one model.
- In single-speaker evaluation with reference speech, it achieved timbre similarity of 0.95 and content error of 0.086, improving over the earlier SwanVoice.
- In instruction evaluations, it ranked 1st in Chinese with APS 86.1 and tied for 1st in English with APS 84.2, and could generate speech, ambient sound, and sound effects in one waveform.
- Weaknesses remain in emotional transitions in complex background music, multi-speaker scenes longer than 2 minutes, continuous control of a particular speaker's emotion, stress, and pauses, and role acting.
Paper links
External research summaries. These are not HDATF publications or measured product results.