AI-Powered Voice Separation Algorithms: Testing Accuracy in Reconstructing Fundamental Frequency for Vocal Analysis.

Journal: Journal of voice : official journal of the Voice Foundation
Published Date:

Abstract

Accurate analysis of the singing voice is often constrained by the need for clean, laboratory-quality recordings. Commercial recordings (CRs)-valuable for artistic and historical study-pose challenges due to complex accompaniment, recording artifacts, and undocumented post-production. This work evaluates whether voice-separation methods can support reliable extraction of fundamental frequency (fo) under such conditions. Synthesized vocals with ground-truth fo were mixed with instrumental introductions from seven CRs of Handel's Ombra mai fu at five signal-to-noise ratios (SNR: -12, -6, 0, +6, +12 dB), yielding 105 samples. Two baselines (unfiltered; bandpass 2.2-6 kHz) and three separation methods (iZotope RX10, Music.ai, robust principal component analysis [RPCA]) were applied. fo contours were extracted in Praat and compared to ground truth using (i) success rate (valid fo detections), (ii) resolved rate (≤10 cents deviation), and (iii) receiver operating characteristic analysis. A clear hierarchy emerged: among the tested methods, Music.ai showed the most robust overall performance, typically exceeding 80% success at SNR ≥ 0 dB and degrading least at low SNR. iZotope RX10 performed similarly at positive SNRs but declined more with noise. Bandpass filtering performed comparably to the best separation techniques at higher SNR levels, whereas RPCA achieved lower accuracy overall. Accuracy decreased below 0 dB across methods and was strongly affected by accompaniment complexity and recording quality. Importantly, with per-segment accuracy near 80%, multiple independent segments are needed to reach high confidence, while correlated samples offer little gain. These results demonstrate that artificial intelligence-based separation methods can open new possibilities for analyzing singing voices in commercial and archival recordings, making large-scale studies of vocal style and technique more feasible than before. At the same time, careful validation across genres, recording conditions, and real performances remains essential to ensure that gains in fo tracking accuracy translate into reliable insights on spectral balance, timbre, and artistic expression. By clarifying both the promise and the boundaries of current approaches, this study provides a foundation for future research that seeks to bridge laboratory analysis and the complexity of real-world recordings.

Authors

Keywords

No keywords available for this article.