Comparative evaluation of handwriting recognition by large language models (LLMs) in interpreting handwritten medication lists.

Journal: Exploratory research in clinical and social pharmacy
Published Date:

Abstract

INTRODUCTION: Handwritten bedside medication lists remain common in healthcare, especially in low-resource countries, presenting challenges for digitization and automated safety checks. The introduction of large-language models may present an opportunity to facilitate digitization of handwritten medication lists. OBJECTIVE: This study evaluated the accuracy of three GPT-based language models in recognizing handwritten medication lists, presented in Dutch. METHODS: Thirty-three participants transcribed a list of 10 medications with dosage and administration instructions. These lists were processed by each model and scored for correct medication name, dosage, frequency, and route. The effect of writer characteristics on LLM performance was assessed using the Mann-Whitney U test (sex, handedness, ink colour) and Kruskal-Wallis H-test (handwriting style: cursive, print, mixed). RESULTS: GPT 4.1 achieved the highest accuracy, followed by GPT4o, both outperforming GPT4o-mini (p < 0.001). Recognition strongly correlated with human legibility (ρ = 0.655; p < 0.001). Print handwriting and blue ink resulted in higher recognition than cursive or mixed styles and black ink. Complex dosing instructions were most error-prone. CONCLUSION: GPT-based optical character recognition showed potential for scalable digitization of handwritten prescriptions, however, human oversight remains essential to ensure medication safety. Future research should validate performance in real-world, multilingual settings.

Authors

Keywords

No keywords available for this article.