Concordance of ChatGPT, Gemini, Claude, and OpenEvidence with the 2024 AAOS guidelines on acute isolated meniscal pathology.
Journal:
The Knee
Published Date:
Mar 21, 2026
Abstract
PURPOSE: To evaluate the reliability and clinical applicability of the three most commonly used large language models (LLMs) (ChatGPT, Gemini, and Claude) and a domain-specific artificial intelligence (AI) platform (OpenEvidence) in providing recommendations for acute isolated meniscal pathology, to compare their accuracy, and to assess the consistency between American Academy of Orthopedic Surgeons (AAOS) Clinical Practice Guidelines (CPG) recommendations and AI-generated guidance. METHODS: An exploratory cross-sectional benchmarking analysis evaluated concordance of three large language models (ChatGPT, Gemini, Claude) and one domain-specific AI (OpenEvidence) with 2024 AAOS clinical practice guidelines for acute isolated meniscal pathology. Nine guideline recommendations were converted into standardized questions and presented to each AI model on the same day. Three sports medicine orthopedic specialists independently assessed responses as concordant or discordant, with disagreements resolved by majority decision. Statistical analysis used SPSS 29, employing Cochran's Q test for concordance assessment and Fleiss' kappa for inter-rater reliability. RESULTS: OpenEvidence achieved perfect concordance (9/9, 100%), followed by ChatGPT (8/9, 89%), Gemini and Claude (both 7/9, 78%). Overall concordance rate was 86% (31/36). Concordance was 100% for strong and consensus recommendations, 75% for moderate and limited recommendations. Cochran's Q test showed no significant difference among models (Q = 3.00, p = 0.392). Inter-rater reliability demonstrated almost perfect agreement (κ = 0.825, 95% CI: 0.637-1.014). CONCLUSIONS: Although ChatGPT, Gemini, and Claude demonstrated high concordance with the AAOS CPG for acute isolated meniscal pathology, their responses were not consistently guideline-concordant. OpenEvidence achieved the highest descriptive concordance rate (100%); however, statistical superiority could not be established due to the limited number of guideline items. This exploratory benchmarking analysis suggests that domain-specific AI models may represent a valuable tool for retrieving information on acute isolated meniscal injuries. CLINICAL RELEVANCE: The difference between domain-specific AI model and general LLMs underscore the need to educate the general public and clinicians about the limitations of general-purpose chatbots, emphasizing that LLM outputs should be interpreted with caution in real-world practice, while tools like OpenEvidence exist for evidence-based information.
Authors
Keywords
No keywords available for this article.