dots.tts Technical Report

Published
Source
arXiv
Paper number
364
Field
AI / General
arXiv ID
2606.07080

Key points

  • It trains a high-resolution AudioVAE that encodes 48 kHz speech into a 25 Hz, 128-dimensional latent space and adds semantic structure through WavLM alignment loss.
  • It uses a three-module split of semantic encoder, LLM, and AR flow-matching head to reduce long-range error accumulation.
  • It introduces SOAR self-corrective alignment, which exposes the flow-matching head to its own inference errors without a reward model.
  • CFG-aware MeanFlow distillation enables 2 to 4 NFE inference steps and achieves 0.231 RTF and 54 ms first-packet latency on H800.
  • It reports open-source SOTA on Seed-TTS-Eval with WER 2.92 and SIM 79.2, and releases the full system under Apache 2.0.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)