Integrity of artificial intelligence and machine learning evidence in spine surgery: a critical appraisal of systematic reviews.

Journal: Journal of neurosurgery. Spine
Published Date:
(1)

Abstract

OBJECTIVE: Artificial intelligence (AI) and machine learning (ML) are increasingly integrated into spine surgery across diagnostic, perioperative, and predictive applications, driven by their potential to enhance accuracy, efficiency, and patient safety. As this literature on AI and ML in spine surgery rapidly expands, concerns have emerged regarding selective reporting and overemphasis on algorithmic performance without adequate attention to transparency, bias, or clinical applicability. Such practices raise the risk of reporting spin, particularly within systematic reviews and meta-analyses, which strongly influences clinical interpretation. This study aimed to evaluate the prevalence and types of reporting bias (spin) in systematic reviews and meta-analyses on AI and ML in spine surgery, and to assess their overall methodological and reporting quality. METHODS: Following PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) guidelines, the authors searched the PubMed, Scopus, and Embase databases using the terms "spine surgery" AND ("artificial intelligence" OR "machine learning") AND ("systematic review" OR "meta-analysis"). Eligible peer-reviewed reviews were evaluated for 15 types of spin, and methodological quality was assessed using AMSTAR 2 (A Measurement Tool to Assess Systematic Reviews 2). Study characteristics included PRISMA adherence, publication year, and level of evidence. Associations with spin prevalence were examined using chi-square, Mann-Whitney U, and Kruskal-Wallis tests and Spearman's rank correlation. RESULTS: The search identified 273 studies; after removing 90 duplicates and 124 ineligible titles/abstracts, an additional 22 studies were excluded to reach the 37 total. Spin was present in 32 of 37 (86.5%) abstracts. All spin types were observed, with the most common being type 3 (18/37, 48.65%), type 9 (18/37, 48.65%), and type 5 (19/37, 51.35%). Lower journal impact factors and reduced AMSTAR 2 confidence ratings were significantly associated with spin (p < 0.05). CONCLUSIONS: Systematic reviews and meta-analyses of AI and ML in spine surgery demonstrate a high prevalence of spin (86.5%), most often overstating conclusions, selectively reporting outcomes, and making inappropriate generalizations. Spin was significantly associated with lower journal impact factors and weaker methodological quality, raising important concerns regarding the reliability and clinical interpretability of the existing evidence base. These findings highlight the need for greater transparency, adherence to reporting standards, and methodological rigor to ensure accurate communication of AI and ML applications. Addressing these biases is critical for reliable evidence synthesis and informed clinical decision-making in spine surgery.

Authors

Keywords

No keywords available for this article.