DisNet : Learning interpretable depression representations in speech.

Journal: Neural networks : the official journal of the International Neural Network Society
Published Date:

Abstract

Speech-based depression detection (SDD) offers an objective and convenient method for depression screening and intervention. However, existing deep learning methods lack interpretability, which prevents them from revealing their findings for the quantitative detection of depression. To address these issues, we propose an interpretable Depression screening Network (DisNet), which contains a Learnable frequency-domain FilterBank (LFB) module and a Hierarchical speech Representation Extraction (HRE) module. LFB utilizes a learnable filterbank to select effective frequency bands in speech for generating depression features, while HRE explores interpretable representations and their patch regions contained in the LFB learned features through incremental sparsification and compression operations. DisNet enables end-to-end interpretability. Furthermore, a self-supervised learning strategy, SLRD, is employed to enhance feature interpretability by aiding DisNet in exploring emotional representations in advance. We collect the AMHS-corpus datasets and analyze them from cross-sectional and longitudinal perspectives. We also evaluate the DisNet's performance on the public datasets DAIC-woz, CMDC and EATD-corpus, yielding average F1 score improvements of 15.9%, 2.8%, and 13%, respectively. Beyond performance, DisNet identifies pronunciation variations of phonemes /i/, /a/, /e/ and /u/ within the depression group and highlights their frequency ranges. In particular, we offer a visualization method as an alternative to conventional scale assessments.

Authors

Keywords

No keywords available for this article.