Diagnostic utility of artificial intelligence in musculoskeletal physical therapy: A comparison with physical therapists.

Journal: Musculoskeletal science & practice
Published Date:

Abstract

BACKGROUND: Although generative artificial intelligence (AI) shows promise in healthcare, its diagnostic performance relative to physical therapists (PTs) in musculoskeletal conditions remains largely unknown. Limited research has examined how generative AI diagnostic reasoning compares to PTs. OBJECTIVES: Generative AI's diagnostic performance in musculoskeletal cases was examined by assessing case accuracy (correct final diagnosis), efficiency (when the correct diagnosis was selected), and funneling behavior (progressive narrowing of diagnoses), compared to PTs of varying expertise. DESIGN: Cross-sectional comparative case-based study evaluating generative AI diagnostic ability. METHODS: Generative AI models (ChatGPT [GPT-4/GPT-4 Turbo] and Gemini [Pro Model and 2.0 Flash]) were evaluated using five standardized musculoskeletal clinical cases presented in a sequential, clue-based format. AI responses were scored for accuracy, efficiency, and funneling behavior. Each AI model completed five clinical cases in 15 independent trials. AI performance was compared with previously published responses from 1201 PTs across experience levels who completed the same cases. RESULTS/FINDINGS: AI responses achieved the highest overall accuracy rates (20.0-83.3%) compared with specialist PTs (19.7-79.6%) and non-specialists (7.6-60.4%) (p < 0.001), with generally higher efficiency. However, specialist PTs outperformed AI in both accuracy and efficiency in 2 of 5 cases, specifically for the lumbar spine and hip (p < 0.001). Specialist PTs also demonstrated greater funneling behavior (up to 50.0%) than generative AI (up to 13.3%; p < 0.001). CONCLUSIONS: Generative AI demonstrated case-dependent agreement with expert-derived diagnoses in standardized musculoskeletal cases, including higher case accuracy and comparable efficiency to specialists. However, the constrained case design limits generalizability and does not reflect open-ended clinical reasoning.

Authors

Keywords

No keywords available for this article.