Not Such a Long Way Off? Contemporary Artificial intelligence Performance Evaluation on Adult Medicine “Long Cases”

Journal: medRxiv
Published Date:

Abstract

The Royal Australasian College of Physicians (RACP) “Long Cases” assess a trainee’s clinical reasoning beyond what is tested in multiple-choice questions, which large language models (LLMs) have already demonstrated proficiency in. This study evaluated a LLM’s ability to perform a “Long Case” assessment, including history-taking, case presentation, and answering examiner questions. The LLM achieved passing consensus scores of 4-5 out of 6 on five cases, suggesting potential for LLMs in complex clinical evaluations.

Authors

  • Christina Gao; Jamie Bellinge; Shaddy El-Masri; Ivana Chim; Michael Sorich; Ishish Seth; James Gorcilov; Matthew Lim; Liam McCoy; Andrew Vanlint; Lauren Lim; Jessica Stranks; Andrew Zannettino; John Maddison; Stephen Bacchi