Artificial Intelligence Large Language Models for Pelvic Floor Rehabilitation: Readabilty, Usability and Consistency.
Journal:
Neurourology and urodynamics
Published Date:
Jul 17, 2026
Abstract
BACKGROUND: Nowadays, artificial intelligence Large Language Models (LLMs) are widely used by patients and physicians alike to investigate various medical topics. Pelvic floor rehabilitation has become a popular subject in recent years. OBJECTIVES: The aim of this study is to assess and compare the usability, readability, and repeatability of three LLMs - ChatGPT, DeepSeek and Gemini - in relation to pelvic floor rehabilitation. METHODS: A total of 35 questions derived from the three most frequently searched Google Trends keywords related to pelvic floor rehabilitation ('pelvic floor dysfunction', 'pelvic floor exercises', and 'pelvic floor physical therapy') were evaluated by two raters. The quality of the responses was assessed using the Brief DISCERN (BD), a validated tool for evaluating the quality of health information. Readability was assessed using the Flesch-Kincaid Reading Ease (FRE), the Flesch-Kincaid Reading Grade Level (FKRGL) and the Simple Measure of Gobbledygook (SMOG) index. Responses were evaluated at two separate time points to assess consistency. RESULTS: Statistically significant differences in the SMOG index were observed among the AI models at the first and second evaluations (p < 0.001), but only at the second evaluation for FKGL (p = 0.006). Significant differences were also observed between LLMs for BD scores for physical therapy section at both first and second evaluations (p < 0.001, p = 0.035 respectively). CONCLUSIONS: ChatGPT, DeepSeek and Gemini provided readable, useful and repeatable answers to questions related to pelvic floor rehabilitation. However, it is important to bear in mind that LLMs are supplementary tools.
Authors
Keywords
No keywords available for this article.