A Deep Learning system for Automated Quality Control of Medical Questionnaire data in Large-Scale Cohort Studies.

Journal: Journal of molecular cell biology
Published Date:

Abstract

Recent advances in natural language processing combined with deep learning have transformed disease prediction, diagnosis, and treatment, yet their application in automated quality control of medical questionnaire data remains largely unexplored. Here, we present an end-to-end system tailored for large-scale cohort interview data. A fine-tuned Whisper model was developed for automatic audio-to-text transcription and optimized for recognizing domain-specific medical terminology and regional dialects, achieving an accuracy of over 91%. Using the DeepSeek framework, we jointly analyzed transcribed texts, interviewer records, and questionnaire content from 522 medical interviews to identify quality issues, achieving the area under the receiver operating characteristic curve (AUC) values of 0.893, 0.953, and 0.973 for omitted question stems, response-record inconsistencies, and insufficient follow-up, respectively. The system achieved high agreement with the senior specialist reference (F1 = 0.843, Kappa = 0.835) and significantly outperformed the junior specialist group (P < 0.01). Evaluation on an independent validation dataset (n = 96), which was collected using a structurally distinct brain health questionnaire, further demonstrated robust performance and generalizability (F1 = 0.831, Kappa = 0.821). Collectively, this framework enhances the efficiency and reliability of medical questionnaire data processing and offers a practical artificial intelligence-driven solution for quality control in global health data, especially in low-resource dialect settings.

Authors

Keywords

No keywords available for this article.