Brain-like dynamics in speech representations can emerge through self-supervised learning

Journal: bioRxiv
Published Date:

Abstract

Speech representations in the human brain do not simply mirror the instantaneous speech signal; rather, they display several properties that are hypothesized to facilitate the integration of speech sounds into words. In particular, neural encodings of speech maintain information that has dissipated from the acoustics, and have also been argued to abstract over variability in how individual speech sounds are produced. Here, we investigate how such characteristics could arise. We introduce a computational framework that uses modern neural network models from speech technology to examine two factors in particular: the learning mechanism and the learning input. We find that self-supervised models trained without lexical or semantic feedback developed temporal dynamics similar to brain representations, regardless of whether they were trained on speech or non-speech audio. In contrast, only models trained on speech learn to abstract over variability due to word position and phonetic context. Overall, our results suggest that domain-general learning mechanisms can lead to several important properties of speech representations, but in some cases require domain-specific input in order to do so.

Authors

  • Liu
  • O. D.; Tang
  • H.; Feldman
  • N. H.; Goldwater
  • S.

Categories