Human-Edited Generative AI-Assisted Multiple-Choice Questions in Postgraduate Family Medicine: Blinded Cross-Sectional Comparative Psychometric Study.
Journal:
JMIR medical education
Published Date:
Aug 10, 2026
Abstract
BACKGROUND: Generative artificial intelligence (GenAI) is increasingly used to draft multiple-choice questions (MCQs) for health professions education, but much evidence concerns raw model outputs, expert ratings, or item difficulty alone. Educators edit GenAI drafts before use, and whether such items are psychometrically ready for postgraduate assessment remains unclear. OBJECTIVE: This study aimed to compare human-edited GenAI-assisted and educator-crafted MCQs for postgraduate Family Medicine Applied Knowledge Test-level assessment, examining difficulty, discrimination, reliability, distractor functioning, and participant perceptions. METHODS: We conducted a blinded cross-sectional, within-participant comparative psychometric evaluation in Singapore. Sixty best-of-five single-best-answer MCQs were evaluated, 30 human-edited GenAI-assisted items and 30 educator-crafted items, topic-matched across postgraduate FM domains and randomized across 2 assessment sets. Eligible participants were postgraduate doctors enrolled in FM residency or postgraduate family medicine programs, preparing for the Applied Knowledge Test, and blinded to item origin; incomplete paired responses were excluded. Outcomes included paired total scores, score correlation and agreement, Kuder-Richardson Formula 20 reliability, item difficulty index, corrected point-biserial discrimination, distractor functioning, and perceived difficulty, clarity, and relevance. Analyses used paired-sample tests, Pearson correlation, Fisher exact tests, and item-level psychometric statistics, with α=.05 and Bonferroni correction within comparison families. RESULTS: Of 74 participants, 73 completed both item sets and were included in the analysis. The final sample comprised 36 graduate diploma in FM trainees, 5 MMed FM trainees, and 32 FM residents. Paired-sample testing showed lower scores on GenAI-assisted than educator-crafted items (mean 19.12, SD 2.83 vs mean 21.10, SD 3.42 out of 30; mean difference -1.97, 95% CI -2.72 to -1.23; P<.001; Cohen d=0.62), indicating that GenAI-assisted items were not easier. Scores were positively correlated (r=0.49, 95% CI 0.30-0.64; P<.001), but Bland-Altman analysis indicated limited agreement. Kuder-Richardson Formula 20 reliability was lower for GenAI-assisted items (0.38 vs 0.60). Mean difficulty index did not differ significantly (0.64 vs 0.70; mean difference -0.07, 95% CI -0.19 to 0.06; P=.29), and more GenAI-assisted items fell within the acceptable difficulty range (18/30, 60.0% vs 13/30, 43.3%). However, mean corrected point-biserial discrimination was lower for GenAI-assisted items (0.09 vs 0.18; mean difference -0.08, 95% CI -0.16 to -0.01; P=.04), and negative discrimination was more common (6/30, 20% vs 3/30, 10%). GenAI-assisted items also had more nonfunctioning and negatively discriminating distractors, although these differences were not statistically significant. Participant ratings of perceived difficulty, clarity, and practice relevance did not differ by origin. CONCLUSIONS: Human-edited GenAI-assisted MCQs can achieve plausible difficulty, but difficulty and surface acceptability did not ensure assessment readiness. Using trainee response data, this study extends work on raw outputs or expert opinion. GenAI should be used as a drafting adjunct within educator-led workflows prioritizing key verification, distractor engineering, pilot testing, empirical item analysis, and repair before item-bank or summative use.
Authors
Keywords
No keywords available for this article.