Challenges in using natural language processing to stratify students by narrative assessments in undergraduate medical education.

Journal: Academic medicine : journal of the Association of American Medical Colleges
Published Date:

Abstract

PROBLEM: Medical educators pay increasing attention to the potential utility of narrative data for assessment, but lack of efficient and standardized ways of interpreting the data have limited its use. Natural language processing (NLP) algorithms could provide new insights for using narrative data within assessment especially identifying at-risk or low-performing students. APPROACH: Assessment data were reviewed from 16 cohorts of medical students from the University of Cincinnati College of Medicine (graduating classes of 2006-2022). A T-score Average (TSA) was calculated for each student based on clerkship assessment data. Narrative data from core clerkship evaluations in responses to the prompt "any opportunities for improvement" were utilized for analysis. The narrative data and calculated TSA were then used to train and test 4 NLP models with the goal of utilizing NLP to identify at-risk students as defined by the bottom 10% average TSA score. OUTCOMES: Based on typical NLP methods, the developed NLP models all performed adequately in identifying student performance with an overall accuracy of 0.8 across all 4 models. However, none of the NLP models was able to identify students within the bottom 10% of performance. During this process, we uncovered the presence of "copy/paste" behavior, a previously undocumented phenomenon within narrative data where preceptors duplicated comments from 1 student to the next. Training NLP models including "copy/paste" comments improved NLP ability to identify students within the bottom and top 10% of performance. NEXT STEPS: NLP models were unable to accurately identify at-risk students. Model accuracy increased with the inclusion of "copy/paste" comments, indicating there may be some discriminatory functionality within aspects of narrative data beyond keywords such as lexical diversity and word quantity. Future work will explore how these other aspects of narrative data associate with performance and utilize more advanced large language models of analysis.Based on typical natural language processing (NLP) methods, the tested NLP models performed adequately in identifying student performance with an overall accuracy of 0.8 across all models. However, none of the NLP models was able to identify students within the bottom 10% of performance, revealing widespread "copy/paste" behavior in evaluations that may limit narrative data's discriminatory utility.

Authors

Keywords

No keywords available for this article.