An expert-based assessment of the quality, reliability, understandability, and readability of ChatGPT responses on cerebral palsy.

Journal: Child's nervous system : ChNS : official journal of the International Society for Pediatric Neurosurgery
Published Date:

Abstract

PURPOSE: In recent years, artificial intelligence-based language models have emerged as a means of rapid access to health-related information. This study aimed to evaluate the quality, reliability, understandability, and readability of ChatGPT's responses to frequently asked questions by families of children with cerebral palsy (CP). METHODS: Responses generated by the free version of ChatGPT to the ten most frequently asked questions posed by families of children with CP were obtained. These responses were evaluated across four dimensions: quality, reliability, understandability, and readability. Content quality was evaluated using the DISCERN instrument. Reliability was assessed by 20 physiotherapists holding MSc or PhD degrees using a 5-point Likert scale. Understandability and actionability were assessed using the Patient Education Materials Assessment Tool for Printable Materials (PEMAT-P), while readability was analyzed using the Flesch-Kincaid Grade Level (FKGL). RESULTS: The median DISCERN score was 45, reflecting average content quality. The median Likert scale score across all questions was 4 out of 5, indicating reliable responses. No significant differences were observed between PhD and MSc expert raters in their Likert scale ratings. The median PEMAT-P was 69.23; however, only 40% of answers exceeded the threshold for understandability, and none met actionability threshold. All responses were written at a level exceeding high school, indicating limited readability for the general public. CONCLUSION: ChatGPT has the potential to provide accurate information to families of children with CP; however, improvements in understandability, actionability, and readability are needed to better support families. Further development is required in order to support decision-making and care processes.

Authors

Keywords

No keywords available for this article.