Speaker Identification Using Voice Quality Features: A Psychoacoustic and Machine Learning Approach.
Journal:
Journal of voice : official journal of the Voice Foundation
Published Date:
Apr 23, 2026
Abstract
OBJECTIVES: This study investigated the discriminative power of psychoacoustically derived voice quality features for closed-set speaker identification. Specifically, we aimed to evaluate the extent to which these features can distinguish individual speakers within a known cohort and to identify the most influential features for differentiating speakers in this context. METHODS: Speech recordings were collected from 120 native Persian speakers, comprising 60 males and 60 females. Utilizing the psychoacoustic model of voice, 26 acoustic features were extracted, including 13 static (base) features along with their corresponding dynamic counterparts, represented by the coefficient of variation (CoV). A Random Forest classifier was employed to evaluate speaker separability within a closed-set framework and to assess the relative contribution of individual features to speaker discrimination. RESULTS: For male speakers, classification accuracy reached 90.15% when both base and dynamic features were combined, surpassing that of models that relied solely on base (89.53%) or dynamic (77.70%) features. In contrast, for female speakers, the combined model achieved an accuracy of 90.74%, compared to 91.24% for base features and 76.52% for dynamic features. Feature importance analysis indicated that for males, the most discriminative features were H1*-H2*, CoVH1*-H2*, and f0. For females, f0, CoVf0, and F3 were found to be the most significant. These results suggest that temporal variability enhances speaker discrimination primarily for male voices, while female speaker identity is more effectively captured through stable spectral and pitch-related features. CONCLUSION: Voice quality features extracted from the psychoacoustic model of voice exhibit significant discriminative capacity across speakers of Persian and can be effectively utilized for closed-set speaker identification. The contribution of dynamic information appears to be sex-dependent, enhancing performance in males but not in females. This indicates that temporal versus stable acoustic cues play differing roles in encoding vocal identity. These findings may inform future research in forensic phonetics, speaker recognition systems, and objective voice-quality assessment.
Authors
Keywords
No keywords available for this article.