Artificial intelligence-based decision support in simulated free flap Re-exploration for head and neck reconstruction. A case-based comparative study.
Journal:
JPRAS open
Published Date:
May 16, 2026
Abstract
Free flap compromise after head and neck reconstruction requires rapid recognition and structured escalation. Large language models have shown potential for text-based clinical support, but their alignment with microsurgical decision-making in flap compromise remains unclear. This simulated case study evaluated three large language models, ChatGPT, Gemini, and Copilot, using 30 standardized head and neck free flap compromise scenarios, including 15 intraoperative and 15 postoperative cases. Each model received identical prompts requesting the suspected diagnosis, likely mechanism of compromise, and recommended management steps. Four consultant microsurgeons independently rated each response using 5-point Likert scales across clinical accuracy, management recommendations, assessment correctness, usefulness of explanation, and overall confidence. Potential for harm was rated as yes or no Ratings were aggregated to case level. Between-model comparisons were performed separately for intraoperative and postoperative scenarios using Friedman tests with Bonferroni-adjusted post hoc Wilcoxon signed-rank tests, while potential for harm was compared using chi-square tests. Clinical accuracy ratings were high across all models in both intraoperative and postoperative scenarios, with no significant between-model differences. In intraoperative cases, usefulness of explanation differed significantly between models, with lower ratings for Copilot than Gemini. In postoperative cases, Gemini received significantly higher ratings for management recommendations and usefulness of explanation than ChatGPT and Copilot, and higher overall confidence than Copilot. Potentially harmful recommendations were uncommon, but differed between models in postoperative scenarios. Interrater agreement was generally higher in intraoperative than postoperative cases. Large language models showed high alignment with expert ratings for recognition of flap compromise in standardized simulated scenarios. However, differences emerged in management planning, explanatory clarity, and confidence, particularly in postoperative cases. These findings support cautious interpretation and do not support independent clinical use in microsurgical care.
Authors
Keywords
No keywords available for this article.