dots.tts Technical Report
- Published
- Source
- arXiv
- Paper number
- 364
- Field
- AI / General
- arXiv ID
- 2606.07080
Key points
- It trains a high-resolution AudioVAE that encodes 48 kHz speech into a 25 Hz, 128-dimensional latent space and adds semantic structure through WavLM alignment loss.
- It uses a three-module split of semantic encoder, LLM, and AR flow-matching head to reduce long-range error accumulation.
- It introduces SOAR self-corrective alignment, which exposes the flow-matching head to its own inference errors without a reward model.
- CFG-aware MeanFlow distillation enables 2 to 4 NFE inference steps and achieves 0.231 RTF and 54 ms first-packet latency on H800.
- It reports open-source SOTA on Seed-TTS-Eval with WER 2.92 and SIM 79.2, and releases the full system under Apache 2.0.
Paper links
External research summaries. These are not HDATF publications or measured product results.