Using articulatory feature detectors in progressive networks for multilingual low-resource phone recognitiona).

Journal: The Journal of the Acoustical Society of America

PMID: 39560422

Abstract

Systems inspired by progressive neural networks, transferring information from end-to-end articulatory feature detectors to similarly structured phone recognizers, are described. These networks, connecting the corresponding recurrent layers of pre-trained feature detector stacks and newly introduced phone recognizer stacks, were trained on data from four Asian languages, with experiments testing the system on those languages and four African languages. Later adjustments of these networks include the use of contrastive predictive coding layers at the inputs to those networks' recurrent portions. Such adjustments allow for performance differences to be attributed to the presence or absence of individual feature detectors (for consonant place/manner and vowel height/backness). Some of these differences manifest after feature-level comparisons of recognizer outputs, as well as through considering variations and ablations in architecture and training setup. These differences encourage further exploration of methods to reduce errors with phones having specific articulatory features as well as further architectural modifications.

Authors

Mahir Morshed

Department of Electrical and Computer Engineering, University of Illinois Urbana-Champaign, Urbana, Illinois 61801, USA.
Mark Hasegawa-Johnson

Department of Electrical and Computer Engineering, University of Illinois Urbana-Champaign, Urbana, Illinois 61801, USA.

Keywords

Humans Language Multilingualism Neural Networks, Computer Phonetics Speech Acoustics Speech Production Measurement Speech Recognition Software

External Resources

View on PubMed Access via DOI PubMed (39560422)

Using articulatory feature detectors in progressive networks for multilingual low-resource phone recognitiona).

Abstract

Authors

Keywords

External Resources

Popular Topics

Recent Journals