Evaluating the diagnostic accuracy of ChatGPT-4 in oral and maxillofacial lesions: a pilot clinical vignette study.

Journal: International journal of oral and maxillofacial surgery
Published Date:

Abstract

There is growing interest in whether artificial intelligence (AI) large language models such as ChatGPT-4 can support clinicians and trainees in the diagnosis of oral and maxillofacial surgery (OMFS) lesions. This pilot study evaluated the diagnostic accuracy of ChatGPT-4 in identifying OMFS lesions using standardized clinical vignettes and assessed the clarity of its diagnostic reasoning compared to expert consensus. Fifty diverse clinical vignettes representing a range of OMFS lesions were developed and validated by three oral and maxillofacial surgeons. Each vignette was entered into ChatGPT-4 with a uniform prompt. The AI's 'most likely diagnosis' was compared to expert consensus. Outcomes included overall and category-wise diagnostic accuracy and expert-rated clarity of diagnostic reasoning. ChatGPT-4 showed overall diagnostic accuracy of 70% (35/50), performing best in odontogenic infections (90%) and worst in soft-tissue malignancies (33%). The model's reasoning clarity received an average score of 3.8 out of 5. While ChatGPT-4 excelled in recognizing classical lesion patterns, it showed limitations in interpreting complex cases. ChatGPT-4 demonstrates moderate diagnostic capability for common OMFS lesions and holds promise as an educational tool. However, its limited performance in complex diagnoses underscores the need for domain-specific optimization and expert oversight before any clinical use.

Authors

Keywords

No keywords available for this article.