Assessment of Clinical Competence Using AI-Standardized Clinical Examination

Journal: medRxiv
Published Date:

Abstract

BACKGROUND: As artificial intelligence (AI) enters clinical practice, concerns about deskilling and never-skilling make the independent assessment of clinical competence increasingly important. The AI-Standardized Clinical Examination (ASCE) uses AI to simulate and score standardized clinical encounters, potentially overcoming some of the logistical and standardization constraints of conventional objective structured clinical examinations (OSCEs). Whether it provides valid and reliable assessment at scale is unknown. METHODS: In this prospective multicenter cohort study, final-year medical students from six French medical schools completed a 10-station ASCE six weeks before the national OSCE session. GPT-4o simulated standardized clinical partners; transcripts were subsequently scored by GPT-4.1 against faculty-defined rubrics. The primary outcome was the correlation between ASCE and national OSCE scores. Secondary outcomes were inter-rater reliability of expert transcript scoring, alignment of ASCE scores with expert ratings, and internal consistency. RESULTS: Among 948 enrolled students, 588 had paired ASCE and national OSCE scores. ASCE scores were correlated with national OSCE performance (Spearman's {rho}, 0.57; 95% CI, 0.51 to 0.63), with less than 1% of residual variance attributable to center (intraclass correlation coefficient, 0.01). Eight experts scored 20 randomly selected ASCE transcripts; Krippendorff's was 0.77 overall, 0.80 for clinical skills, and 0.70 for communication and attitude skills. Adding GPT-4.1 as an evaluator reduced by 0.02, and its scores did not differ significantly from those of individual experts. Internal consistency was 0.67 across 10 stations and was projected to approach 0.80 with 20 stations. CONCLUSIONS: Scores from an examination delivered and scored by AI were associated with performance on a national high-stakes OSCE, with minimal center-level variation, tentative agreement among expert raters, and automated scores that aligned with expert ratings. These results support AI delivery and scoring of clinical examinations, with expert review of automated scores.

Authors

  • Lopez
  • A.; Gabellier
  • L.; Peytavi
  • C.; Loi
  • Z.; Poupin-Lavigne
  • E.; Plotton
  • C.; Dupont
  • V.; Jahdauti
  • L.; de Vries
  • P.; Thuillier
  • P.; Compagnat
  • M.; Vergne-Salle
  • P.; Kuchenbuch
  • M.; Feigerlova
  • E.; Berthelot
  • P.; Bednarek
  • N.; Cochener-Lamard
  • B.; Robert
  • P.-Y.; Braun
  • M.; Morin
  • D.; Deruelle
  • P.; Laffont
  • I.; Pers
  • Y.-M.; Yauy
  • K.