Performance of generative artificial intelligence in oral and maxillofacial radiology based on the board-certification examination of Japan.
Journal:
Oral radiology
Published Date:
Aug 20, 2026
Abstract
OBJECTIVE: To evaluate the performance and potential utility of generative artificial intelligence (AI) in oral and maxillofacial radiology using the board-certification examination administered by the Japanese Society for Oral and Maxillofacial Radiology (JSOMR). METHODS: The responses generated by ChatGPT for multiple-choice questions from the board-certification examination of the JSOMR over the three-year period from 2020 to 2022 were assessed. The questions were manually entered individually as prompts for GPT-3.5, GPT-4, and GPT-5, which are the models available from ChatGPT. The accuracy was calculated according to examination year, question format, and level of taxonomy. RESULTS: GPT-3.5 achieved an accuracy of 40.3% for the three years (42.9%, 42.0%, and 36.0% for 2020, 2021, and 2022, respectively), that of GPT-4 was 67.8% (67.3%, 74.0%, and 62.0%, respectively), and that of GPT-5 was 76.5% (79.6%, 78.0%, and 72.0%, respectively). GPT-5's results exceeded the passing score for each year with the accuracy significantly outperformed that of GPT-3.5 and GPT-4. Regarding performance according to the question format, GPT-5 performed significantly superior to the earlier models, especially on two-answer questions. CONCLUSIONS: On the board-certification examination of the JSOMR, the performance of GPT-5 was significantly superior to that of GPT-3.5 and GPT-4. This suggests that, given the rapid development of generative AI, GPT-5 has reached a level of text-based knowledge equivalent to that assessed in the board-certification examination of the JSOMR.
Authors
Keywords
No keywords available for this article.