Automated Generation and Human Evaluation of Neurosurgical Board Examination Self-Assessment Questions.

Journal: Neurosurgery practice
Published Date:

Abstract

BACKGROUND AND OBJECTIVES: Multiple-choice questions are the primary assessment format for neurosurgical board certification. Creating high-quality examination questions requires significant expert time and resources. The goal of this study was to develop an automated system to generate board-style neurosurgical multiple-choice questions using state-of-the-art vision-language models and compare their quality with authentic self-assessment questions. METHODS: We developed an automated pipeline using OpenAI generative pre-trained transformer (GPT)-4o and Anthropic Claude Sonnet-3.5 to generate neurosurgical board-style questions from Neurosurgery Publications articles. We generated 89 587 synthetic questions: 45 689 with GPT-4o and 43 898 with Claude. Each question was associated with a single image extracted from the articles' figures. We evaluated the quality of synthetic questions through 5 surveys comparing 20 synthetic questions (10 from each model) with 10 authentic questions from the Self-Assessment for Neurological Surgeons (SANS) question bank. Each survey was completed by a neurosurgery resident and an attending who guessed the source [human vs artificial intelligence (AI)-generated] and rated suitability for board examination use. We also evaluated the question-answering performance of the generalist GPT-4o and the specialized CNS-Obsidian. RESULTS: SANS questions were more often perceived as human-made than GPT-generated (residents, P = .0002; attendings, P = .1091) and Claude-generated (residents, P = .0002; attendings, P = .0272) questions. Notably, 54% of AI-generated questions misled at least one evaluator, and 23% misled both. In quality assessments, SANS questions outperformed GPT-generated (residents, P < 10-5; attendings, P = .0001) and Claude-generated (residents and attendings, P < 10-5) questions. Particularly, 25% of AI-generated questions were rated as suitable for board examinations vs 72% of human-generated questions when measured by evaluator consensus (P < 10-7). CONCLUSION: Although quality gaps exist between AI-generated and human-created neurosurgical board examination questions, our approach demonstrates the potential of vision-language models to augment assessment development in specialized medical fields, reducing the burden on examination boards and credentialing organizations.

Authors

Keywords

No keywords available for this article.