Comparing Artificial Intelligence versus Human Screening in Systematic Reviews
Journal:
medRxiv
Published Date:
Jul 2, 2026
Abstract
Introduction. Systematic reviews are essential for informing health policy and practice. Artificial intelligence (AI) automates the article screening process and produces time savings, although the performance of AI screening compared to traditional human screening remains uncertain. We undertook this study to compare the performance of two agentic AI tools, namely Loon LensTM and Catchii, to one another and to humans at the title and abstract screening level. We also compared Loon Lens to humans at the full-text screening level. Methods. We developed a de novo research question on the association between any of three ambient air pollutants (carbon monoxide, ozone, nitrogen dioxide) and the onset or worsening of Parkinson disease. A health sciences librarian developed the literature search strategy and we proceded with human screening guided by PRISMA. We uploaded the retrieved citations and the eligibility criteria to both AI tools and compared screening results using sensitivity, specificity, positive predictive value, negative predictive value, concordance, kappa, and F1 score. We compared the calculated performance statistics to those obtained by naive guessing and regressed concordance (agree or disagree with the human reference standard) onto confidence scores provided by Loon Lens, which assigned a confidence level (Very High, High, Medium, or Low) to each of its screening decisions. Human screening was the reference standard against both AI tools and Catchii was the reference standard against Loon Lens. Results. At title and abstract screening, Loon Lens outperformed Catchii when humans were the reference standard. At full-text screening, most disagreements centered around articles Loon Lens included and humans excluded. At both screening levels, higher confidence scores were associated with lower odds of disagreement between Loon Lens and human screeners. Discussion. Given the panoply of available AI screening tools and their differential performance, plus the rapidly evolving nature of AI technology, researchers should pilot test their chosen tool at the start of each review. Sensitivity, kappa, and F1 are the optimal performance statistics to employ, especially at title and abstract screening, where the imbalance between proportions of included and excluded citations can inflate concordance and negative predictive value.