On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin
- Published
- Source
- arXiv
- Paper number
- 889
- Field
- Research
- arXiv ID
- 2608.11898
Key points
- The full training set is about 100,000 hours of standard Mandarin and dialect speech, and the final self-distillation stage uses about 5,000 hours of dialect speech with high error rates.
- The student uses its own generated prefix as input, while the teacher also consults the ground-truth sentence and provides probabilities for each next-output candidate; only the student model is used in deployment.
- Evaluation is performed on eight standard-Mandarin public datasets, five public dialect datasets, and 18 internal dialect datasets.
- Compared with the base model, the average character error rate drops from 3.46 percent to 3.27 percent for standard Mandarin, from 15.37 percent to 12.79 percent for public dialects, and from 21.01 percent to 12.42 percent for internal dialects.
- Continuing ordinary supervised learning on the same data raises the standard-Mandarin error rate to 4.43 percent, while self-distillation keeps it at 3.27 percent, and the overall average error rate is 10.12 percent instead of 10.74 percent.
Paper links
External research summaries. These are not HDATF publications or measured product results.