StepAudio 2.5 Technical Report
- Published
- Source
- arXiv
- Paper number
- 227
- Field
- Research
- arXiv ID
- 2605.23463
Key points
- ASR accuracy: The model achieved an average character error rate (CER) of 2.97% on Chinese benchmarks and a word error rate (WER) of 3.68% on English benchmarks. Its 32K context window makes it especially strong on long-form audio.
- Decoding efficiency: Thanks to the MTP-5 mechanism, the model achieved a real-time factor (RTF) of 0.0053, making it faster than several specialized ASR systems. The acceptance rate of MTP tokens shows a consistent decay factor of about 0.9, validating the choice of five preceding branches.
- TTS preference: In arena-style evaluation, StepAudio 2.5 TTS achieved an overall win rate of 67.6% against existing competitors such as ElevenLabs and Gemini Flash TTS.
Paper links
External research summaries. These are not HDATF publications or measured product results.