StepAudio 2.5 Technical Report

Published
Source
arXiv
Paper number
227
Field
Research
arXiv ID
2605.23463

Key points

  • ASR accuracy: The model achieved an average character error rate (CER) of 2.97% on Chinese benchmarks and a word error rate (WER) of 3.68% on English benchmarks. Its 32K context window makes it especially strong on long-form audio.
  • Decoding efficiency: Thanks to the MTP-5 mechanism, the model achieved a real-time factor (RTF) of 0.0053, making it faster than several specialized ASR systems. The acceptance rate of MTP tokens shows a consistent decay factor of about 0.9, validating the choice of five preceding branches.
  • TTS preference: In arena-style evaluation, StepAudio 2.5 TTS achieved an overall win rate of 67.6% against existing competitors such as ElevenLabs and Gemini Flash TTS.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)