Artificial intelligence versus human consensus: A concordance analysis in the screening of studies for evidence synthesis in physical activity and sport.

Journal: Journal of bodywork and movement therapies
Published Date:

Abstract

OBJECTIVE: To assess the agreement between the ChatGPT Plus (GPT 4.1) version and human consensus during the screening of studies in four different evidence synthesis projects. METHODS: A comparative design was used to analyze the degree of agreement between ChatGPT Plus (GPT 4.1) and human reviewers in the study selection process for two systematic reviews with meta-analyses, one scoping review, and one literature review. Human screening was performed independently using the Rayyan platform, while the artificial intelligence was provided with predefined eligibility criteria and protocols. Screening decisions were compared using Cohen's kappa coefficient, sensitivity, and specificity, using Stata 18. RESULTS: In the systematic reviews with meta-analyses (SR1 and SR2), agreement was high (κ = 0.73 and 0.86), with sensitivity ≥0.88 and specificity ≥0.99, indicating high reliability in excluding irrelevant studies. In contrast, in the scoping review (SR3) and the literature review (NR1), agreement was moderate (κ = 0.56 and 0.59), with lower positive predictive values (≤0.52), suggesting a higher risk of overdetection. CONCLUSION: Overall, the results suggest that AI-assisted screening may serve as a reliable support tool or triage aid in reviews with well-defined inclusion criteria, rather than fully replacing manual screening. However, in more exploratory or interpretive contexts, human oversight remains necessary to ensure the accuracy of the selection process.

Authors

Keywords

No keywords available for this article.