SN-WER: Script-Normalized WER for Multi-Script Indic ASR Evaluation

Published
Source
arXiv
Paper number
303
Field
LLMs / NLP
arXiv ID
2606.02548

Key points

  • Before computing WER, the paper proposes Script-Normalized WER (SN-WER), a training-free, evaluation-only metric that transliterates both reference and hypothesis text into the standard script for each language.
  • In controlled stress tests, artificially induced WER inflation from romanization was reduced by 67%, and in lexical substitution control experiments sensitivity to semantic errors remained almost unchanged, with Delta SN-WER to Delta WER at about 1.09.
  • SN-WER is robust to transliteration choices and normalization changes, and it showed a low token collision rate of under 0.1% in the Indian language settings evaluated.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)