Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding

Published
Source
arXiv
Paper number
933
Field
Research
arXiv ID
2608.16360

Key points

  • The SpREAD dataset is EEG from one participant who listened to 1,353 Japanese sentences spoken by 18 speakers. Each sentence was measured three times on different days, for 45 sessions over 9 days.
  • The train, development, and test sets contain 3,195, 432, and 432 examples, respectively. Repeated measurements of the same sentence were always placed in the same split.
  • On the 432 test examples, the character error rate of 0.948 is significantly lower than the baseline 0.968. A classifier trained on the training set and evaluated on the development set reaches only 2.08%, which is no better than chance at 2.2%.
  • Using contrastive learning alone makes the acoustic correlation coefficient worse, dropping from 0.282 to 0.262. Adding variational regularization recovers it to 0.274, but that is still not better than the baseline, and speaker information improves only marginally.
  • There is only one participant, and the 18 speakers appear in both training and evaluation. The system handles only Japanese speech, and current EEG-based speech reconstruction remains far from practical.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)