A generalizable speech neuroprosthesis

Journal: bioRxiv
Published Date:

Abstract

Intracortical brain-computer interfaces (BCIs) can restore communication to people with vocal tract paralysis by decoding cortical activity during attempted speech into text. State-of-the-art systems pairing neural-to-phoneme decoders with phoneme-to-word language models have achieved word error rates (WERs) as low as 1%, but only after collecting thousands of sentences of training data. Shortening the data collection process would facilitate scaling this new technology by reducing the time from device implant to high-accuracy communication. Here we introduce a transformer-based decoder model trained jointly across six intracortical speech BCI participants. For every participant - regardless of sex, disease etiology, or attempted speaking strategy - a multi-user model decoded speech more accurately (over 50% lower relative WER on average) than models trained on individual users' data. Notably, the multi-user model could be fine-tuned on fewer than 200 sentences from a held-out user to achieve a WER below 7%. These results reveal how to pool intracortical data across people to yield more accurate, generalizable, and rapidly-deployable decoding models.

Authors

  • Fogg
  • Z. M.; Card
  • N. S.; Wairagkar
  • M.; Srinivasan
  • A.; Singer-Clark
  • T.; Hou
  • X.; Okorokova
  • E.; Peracha
  • H.; Iacobacci
  • C.; Brailow
  • T.; Jude
  • J. J.; Levi-Aharoni
  • H.; Le
  • T.; Mifsud
  • D.; Deevi
  • P.; Nason-Tomaszewski
  • S.; Pritchard
  • A. L.; Zhang
  • Y.; Richards
  • B.; Bechefsky
  • P.; Hochberg
  • L. R.; Williams
  • Z.; Shahlaie
  • K.; Au Yong
  • N.; Rubin
  • D.; Pandarinath
  • C.; Brandman
  • D. M.; Stavisky
  • S. D.

Categories