Supporting patient understanding of cervical and ovarian cancer: how well does AI perform?

Journal: Proceedings (Baylor University. Medical Center)
Published Date:

Abstract

BACKGROUND: As the public increasingly interacts with artificial intelligence (AI) chatbots, we compared the answers from five AI chatbots to standardized questions that patients might ask about ovarian and cervical cancer. METHODS: ChatGPT 3.5, Google Gemini 2.0, Reddit Answers, Bootcamp, and DeepSeek were queried with 15 frequently asked questions (FAQs) on cervical cancer and 11 on ovarian cancer. In a blinded, randomized survey, each deidentified response was independently assessed by three gynecologic oncologists using a 4-point scale (1, accurate and comprehensive; 2, accurate but inadequate; 3, accurate but outdated or inaccurate; 4, completely inaccurate). Readability (Flesch-Kincaid grade level) and word count were recorded. RESULTS: For cervical cancer FAQs, ChatGPT 3.5, Google Gemini 2.0, Reddit Answers, Bootcamp, and DeepSeek received scores of 1.4, 1.6, 2.7, 1.5, and 1.5, respectively. For ovarian cancer FAQs, average scores were 1.3, 1.4, 2.5, 1.2, and 1.4, respectively. All AI chatbot responses were written at a reading level above 11th grade, making them generally difficult for the average American to read. CONCLUSIONS: Although generally rated as accurate and adequate, the answers were frequently off-topic and not generalizable. Healthcare providers should be aware of unintentionally generated misinformation to better counsel patients.

Authors

Keywords

No keywords available for this article.