Artificial intelligence responses to toilet training questions: A reliability and readability analysis of ChatGPT and gemini applications.

Journal: Work (Reading, Mass.)
Published Date:

Abstract

BackgroundToilet training is a critical developmental milestone that may have long-term implications for children's psychosocial development if improperly guided. With the increasing use of artificial intelligence (AI) chatbot applications for health-related information, evaluating the reliability and readability of AI-generated guidance on sensitive developmental topics has become essential.ObjectiveThis study aimed to evaluate the reliability and readability of responses provided by AI chatbot applications regarding commonly asked questions about toilet training.MethodsTwo widely used AI platforms-ChatGPT-4 Turbo (OpenAI) and Gemini 2.0 Flash (Google)-were included in the study. The study was initiated on April 29, 2025, by submitting a standardized prompt to the chatbots. Ten frequently asked questions about toilet training were selected based on AI-generated query lists and Google Trends data. Responses were obtained in independent sessions and evaluated by a panel of five child development experts using a four-point Likert-type scale developed by Mika et al. Readability levels were assessed using the Flesch-Kincaid Grade Level through WordCalc software. Statistical analyses were conducted to compare quality and readability across platforms.ResultsStatistically significant differences in response quality were identified for Questions 2, 3, and 7 (p < 0.05). In the quality rating system, lower scores indicate higher response quality. Gemini demonstrated lower (better) median quality scores for these items. Regarding readability, Gemini produced responses with a higher Flesch-Kincaid Grade Level (i.e., more complex reading level), particularly for Questions 3 and 7. No statistically significant differences in response quality were found for the remaining seven questions.ConclusionsBoth AI applications provided generally acceptable expert-rated responses to common toilet-training questions; however, differences in response quality and readability were observed across specific items. These findings suggest that AI tools may serve as accessible supplementary informational resources for families, but they should not be interpreted as substitutes for professional guidance or as evidence of clinical validity.

Authors

Keywords

No keywords available for this article.