Rehab or not rehab: artificial intelligence helps humans identify rehabilitation Cochrane systematic reviews.

Journal: European journal of physical and rehabilitation medicine
Published Date:

Abstract

BACKGROUND: Rehabilitation is a core component of health systems worldwide, yet its conceptual boundaries remain heterogeneous and difficult to operationalize in research and evidence synthesis. Different classification systems have been proposed, including the pragmatic framework introduced by Levack et al. in 2019 and the more structured research-oriented definition developed by Cochrane Rehabilitation in 2022. In parallel, large language models such as ChatGPT have demonstrated increasing potential in automated literature analysis. However, their ability to apply complex domain-specific classification systems has not been systematically evaluated. The aim is to compare the performance of ChatGPT-4.5, expert reviewers, and hybrid human-AI workflows in classifying research papers as rehabilitation or non-rehabilitation using two different classification systems. METHODS: A retrospective inter-rater reliability and diagnostic performance study was performed using online evidence synthesis and Cochrane Library. The sample consisted of 152 Cochrane systematic reviews published between November 2024 and May 2025. Each review was classified as rehabilitation or non-rehabilitation using the Levack framework and the Cochrane Rehabilitation definition. Three evaluation approaches were compared: AI-only classification (using ChatGPT), expert reviewer consensus (gold standard), and a hybrid human-AI workflow with randomized decision order (AI-first or human-first). Agreement with the expert consensus was assessed using Cohen's kappa (k). Diagnostic performance metrics and potential cognitive biases (automation and conservatism bias) in AI-assisted decisions were also analyzed. RESULTS: AI-only classification showed moderate agreement with expert consensus when using the Levack framework (k=0.60, 95% CI 0.47-0.71) and substantial agreement when using the Cochrane Rehabilitation definition (k=0.82, 95% CI 0.69-0.92). Hybrid workflows consistently achieved near-perfect agreement (k-range: 0.97-1.00). AI-only classification demonstrated high negative predictive values but lower positive predictive values, indicating a tendency toward overclassification of borderline cases. No evidence of automation bias or conservatism bias was observed in human-AI collaboration. CONCLUSIONS: Large language models demonstrate meaningful capability in classifying rehabilitation research, although performance varies with the conceptual structure of the applied classification system. Hybrid human-AI workflows achieved the highest accuracy, suggesting that AI is most effective when used as a decision-support tool rather than as an autonomous classifier. Integrating hybrid human-AI workflows into evidence synthesis can significantly accelerate the identification and mapping of rehabilitation literature, ensuring high methodological accuracy without introducing cognitive automation biases for researchers.

Authors

Keywords

No keywords available for this article.