Guideline-based evaluation of large language models for patient counselling before robot-assisted radical prostatectomy: accuracy, safety, and readability.
Journal:
Journal of robotic surgery
Published Date:
Aug 26, 2026
Abstract
Large language models (LLMs) are increasingly used by patients seeking information about prostate cancer surgery, yet their suitability for preoperative counselling remains uncertain. We compared five contemporary LLMs in answering 20 expert-developed questions concerning robot-assisted radical prostatectomy (RARP). First-generation responses were independently assessed by blinded urologists for factual accuracy, information quality, clinical safety and readability using a prespecified reference-answer matrix, DISCERN, EQIP-36, the Global Quality Score, JAMA benchmark criteria and six readability indices. Significant differences were observed across models for all five quality outcomes (all Friedman Pā<ā0.001). ChatGPT achieved the highest median factual accuracy (89.4%), DISCERN score (60), EQIP-36 score (82.0%) and Global Quality Score [5], although none of these outcomes differed significantly between ChatGPT and DeepSeek after multiplicity adjustment. All models exceeded the recommended sixth-grade reading level. ChatGPT generated the most readable responses, with a mean Flesch-Kincaid Grade Level of 9.80 and Flesch Reading Ease Score of 50.27. No potentially harmful responses were identified for ChatGPT; citation coverage was highest for DeepSeek, while citation validity was 94% for both models. Contemporary LLMs can generate generally accurate information about RARP, but model-dependent variability, excessive linguistic complexity and incomplete source transparency limit their unsupervised use. These systems may supplement, but should not replace, individualized counselling by urologists.
Authors
Keywords
No keywords available for this article.