Agreement between expert gynecologists and ChatGPT in the management of suspected retained products of conception.
Journal:
European journal of obstetrics, gynecology, and reproductive biology
Published Date:
Jun 3, 2026
Abstract
OBJECTIVE: To assess agreement between a large language model (ChatGPT version 5.2) and expert gynecologists in the management of suspected retained products of conception (RPOC). METHODS: This case-based comparative agreement study included prospectively collected clinical cases from a tertiary academic medical center. Fifty-five anonymized cases of suspected RPOC were independently evaluated by three senior gynecologists using a predefined management classification system. Cases were categorized as full consensus (agreement among all three experts) or partial consensus (agreement among two experts); cases without consensus were excluded from the primary analysis. For each included case, the consensus-based expert decision served as the reference standard. ChatGPT-generated recommendations were compared with expert consensus under two conditions: blind prediction and a condition with prior exposure to expert decisions. Agreement was assessed using percent agreement and Cohen's kappa (κ). RESULTS: Of 55 cases, 50 met inclusion criteria (24 full consensus; 26 partial consensus). Mean physician agreement with the consensus reference standard was 82.7% (κ = 0.58). ChatGPT demonstrated lower concordance, with accuracy of 34.0% (κ = 0.27) under blind prediction and 54.0% (κ = 0.33) with prior exposure to expert decisions. Performance improved in full-consensus cases (58.3% to 75.0%) but remained substantially lower in partial-consensus cases (11.5% to 34.6%). CONCLUSION: In the management of suspected RPOC, ChatGPT (version 5.2) may currently be more appropriate as a supervised adjunct than as an independent clinical decision-maker.
Authors
Keywords
No keywords available for this article.