Quality of Tinnitus Information From Generative AI Systems and Web Search: An Expert‑Rated Comparative Study.

Journal: Otology & neurotology : official publication of the American Otological Society, American Neurotology Society [and] European Academy of Otology and Neurotology
Published Date:

Abstract

OBJECTIVE: To compare the quality of tinnitus-related information generated by multiple generative artificial intelligence (GenAI) systems and web search using expert evaluation. STUDY DESIGN: Cross-sectional comparative study. SETTING: Digital platforms evaluated in their native public interfaces. PATIENTS: Not applicable. Thirty commonly searched tinnitus-related questions derived from United States Google Trends data (2020-2025). INTERVENTIONS: Questions were submitted to 6 GenAI systems (OpenEvidence, Claude, DeepSeek, GPT-5, Gemini, GPT-4) and Google Search (first organic result). Responses were independently rated by 6 experts using the QAMAI framework. MAIN OUTCOME MEASURES: Mean expert-rated quality scores across 5 domains (accuracy, clarity, relevance, completeness, and usefulness). RESULTS: Overall quality differed significantly across systems (Friedman P<0.001; Kendall W=0.34). OpenEvidence achieved the highest mean score (4.45±0.72; 95% CI: 4.40-4.49), followed by Claude (4.00±1.02), DeepSeek (3.92±1.13), GPT-4 (3.89±0.84), Gemini (3.62±0.98), and GPT-5 (3.30±1.11). Google Search scored lowest (2.27±1.12; 95% CI: 2.20-2.35). Completeness was the lowest-performing domain across systems (range: 1.70-4.41). Pairwise comparisons showed significant differences between OpenEvidence and all other systems (effect size r=0.49-0.86). Inter-rater reliability was high (ICC=0.82). Readability demonstrated an inverse pattern relative to expert-rated quality. OpenEvidence demonstrated the lowest readability (Flesch-Kincaid Grade Level 17.5; Flesch Reading Ease 2.2), corresponding to a postgraduate reading level, whereas general-purpose LLMs produced more accessible responses at a sixth-seventh grade reading level. CONCLUSIONS: The quality of tinnitus information varies substantially across digital platforms. While GenAI systems generally outperform web search, deficiencies in completeness persist. Readability analysis revealed an inverse relationship between expert-rated quality and response accessibility, suggesting that clinician and patient assessments of informational value may not always align. These findings highlight the need for continued evaluation and clinician oversight to ensure safe, comprehensive, and accessible patient-facing information.

Authors

Keywords

No keywords available for this article.