[Would artificial intelligence pass a Geriatrics exam?: evaluating conversational responses and their accuracy].

Journal: Revista espanola de geriatria y gerontologia
Published Date:

Abstract

OBJECTIVE: To evaluate the accuracy of six different artificial intelligence (AI) models in answering questions from Spain's MIR (Médico Interno Residente) exam related to Geriatrics. METHODS: The performance of six AI models was analyzed by comparing their conversational responses with the official answer keys of the MIR exam. The accuracy of each model was measured based on the number of correct answers obtained for the Geriatrics questions. Due to the sample size, a descriptive analysis of the accuracy percentages was performed. The questions were grouped into specific thematic categories, and the linguistic complexity of the justified responses was analyzed using the formal application of the INFLESZ readability index. RESULTS: The models demonstrated a high accuracy rate: peak performance reached 97.22%, an intermediate group achieved 91.67%, and the minimum score was 83.33%. Only one clinical question (a 90-year-old patient with insomnia due to gonalgia secondary to gonarthrosis) was failed simultaneously by five of the six analyzed tools. The overall INFLESZ index ranged between 44.63 and 49.26, classifying the explanatory texts as «somewhat difficult». CONCLUSIONS: AI models constitute tools of remarkable reliability and potential utility for theoretical medical education environments due to their high correct answer rates. However, they exhibit evident limitations in real-world clinical practice owing to the variability of individual medical scenarios, which are not always reproducible through traditional algorithmic design. Therefore, their role in current medicine must be conceived as complementary and never as an autonomous system for diagnosis or healthcare decision-making.

Authors

Keywords

No keywords available for this article.