Scalable Keyword Spotting via Modular Network Expansion

Published
Source
arXiv
Paper number
697
Field
Research
arXiv ID
2607.19918

Key points

  • It freezes the entire base network, including batch-normalization statistics, so existing keyword performance is preserved at the pixel level.
  • With only 10K additional parameters, 6.7% of the 150K base, it improves FRR for new keywords from 6.46% to 4.37%.
  • It achieves better performance with fewer MACs, 16.34M versus 18.45M and 20.52M, than Adapter or LoRA.
  • A core-first decision rule preserves the exact behavior of the existing keyword threshold.
  • It can scale using only new keyword data, without access to the original training data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)