Simulation-based evaluation of ChatGPT for healthcare associated infection surveillance using validated case scenarios.

Journal: American journal of infection control
Published Date:

Abstract

BACKGROUND: Surveillance for healthcare-associated infections is central to infection prevention but remains complex, resource-intensive, and variable. Large language models like ChatGPT offer potential support but have not been evaluated for applying National Healthcare Safety Network (NHSN) definitions. METHODS: This cross-sectional simulation assessed ChatGPT (GPT-4) accuracy and consistency in applying NHSN definitions. A 20-item test from a validated training bank was categorized by domain. Three independent runs were conducted using identical prompts and the 2024 NHSN manual, accessed through the ChatGPT Plus platform. ChatGPT's responses were compared to a validated answer key and benchmarked against infection preventionists (n = 22). Analyses included Wilcoxon signed-rank, Welch's t-test, and Spearman correlation. RESULTS: ChatGPT averaged 45% accuracy across runs, with only 20% of items correct in all 3 attempts. It was more reliable on laboratory-identified events but inconsistent on infections requiring temporal or multi-step logic. No significant correlation emerged between ChatGPT correctness and IP accuracy. CONCLUSIONS: In this simulation-based proof-of-concept, ChatGPT-4 demonstrated limited reliability for complex NHSN surveillance tasks but showed potential in structured, lab-based scenarios. These findings highlight both the current constraints and future opportunities for generative artificial intelligence tools in infection prevention.

Authors

Keywords

No keywords available for this article.