Large Language Models for Ophthalmology Training in China: A Prospective Evaluation.

Journal: Ophthalmology science
Published Date:

Abstract

PURPOSE: This study explored large language models (LLMs) as a scalable solution to the global shortage and uneven distribution of ophthalmologists, particularly their actual effectiveness and potential risks in ophthalmic training. DESIGN: This is a prospective study. SUBJECTS: Eleven leading LLMs (ERNIE Bot 4.5 Turbo, Dou Bao, Tongyi Qianwen 3, DeepSeek-R1, ChatGPT-4o, Gemini 2.5 Flash, Tencent Yuanbao T1, HuatuoGPT II, Baichuan4-Turbo, Kimi k1.5, GLM-4) and 10 resident physicians (RPs) from a tertiary ophthalmology hospital in China. METHODS: Phase 1: all LLMs were tested on the Chinese and English versions of the Chinese National Health Professional Technical Qualification Examination (Intermediate Level) in Ophthalmology (CNHPTQE-O). Phase 2: the best-performing LLM was used to assist the 10 RPs in 2 tasks: (1) answering the same CNHPTQE-O text questions and (2) classifying 4 types of keratitis images. Resident physicians completed unassisted and assisted phases with a 1-month washout period. MAIN OUTCOME MEASURES: The primary outcomes were accuracy (%) on the CNHPTQE-O for LLMs, and change in accuracy for RPs with versus without LLM assistance (text and image tasks). The secondary outcomes included postassessment survey ratings and confusion matrix analysis. RESULTS: Several Chinese LLMs, especially ERNIE Bot 4.5 Turbo, demonstrated superior performance on the CNHPTQE-O, achieving accuracies of 98.00% (Chinese) and 86.50% (English). ERNIE Bot 4.5 Turbo significantly outperformed all RPs on the Chinese examination (P = 0.001). With LLM assistance, all 10 RPs passed the text examination; the mean accuracy improved from 60.75% to 79.00% (mean difference 18.25%, 95% confidence interval: 10.30%-26.20%, P = 0.001). Questionnaire feedback was positive. However, on the keratitis image task, LLM assistance did not improve RP accuracy (41.25% vs. 40.56%, P = 0.662); questionnaire feedback was markedly less favorable. CONCLUSIONS: Large language models possess a solid foundation in ophthalmic knowledge and can effectively enhance trainee performance in text-based assessments, demonstrating clear potential as a training aid. However, their limitations in image-assisted diagnostic tasks and the associated risk of "artificial ignorance" should not be overlooked. FINANCIAL DISCLOSURES: The author has no/the authors have no proprietary or commercial interest in any materials discussed in this article.

Authors

Keywords

No keywords available for this article.