Transforming Systematic Reviews: Evaluating a Fine-Tuned Large Language Model for Abstract Screening in Uveitis and Retinal Vasculitis: Fine-Tuned LLM for Review Screening.

Journal: Ophthalmology science
Published Date:

Abstract

PURPOSE: To evaluate the classification performance of UveAItis, a domain-specific large language model (LLM) fine-tuned for automated title and abstract screening in systematic reviews, using retinal vasculitis as a prototype. DESIGN: Comparative evaluation study embedded within a registered systematic review and meta-analysis (PROSPERO: CRD42023489232). SUBJECTS: A total of 1030 randomly selected articles from an initial search of 5533 records related to retinal vasculitis. METHODS: Articles were independently screened by 2 uveitis experts (gold standard), final-year medical students, and 3 LLMs: UveAItis (fine-tuned Generative Pre-trained Transformer [GPT]-4o), base GPT-4o, and Claude Sonnet 3.5. Screening followed a 2-question binary logic regarding human subjects and primary empirical research design. Discrepancies were resolved through expert adjudication. MAIN OUTCOME MEASURES: Classification accuracy, sensitivity, specificity, area under the receiver operating characteristic curve, and Cohen Kappa coefficient for inter-rater agreement. RESULTS: UveAItis achieved the highest performance with an accuracy of 93.3%, area under the curve (AUC) of 0.887, and Kappa of 0.77. It significantly outperformed base GPT-4o (AUC: 0.805, P = 0.021), Claude Sonnet 3.5 (AUC: 0.669, P < 0.0001), and medical students (AUC: 0.585, P < 0.00001). The fine-tuned model correctly identified 65.4% of expert-included articles postconsensus, whereas students only identified 22.3%. UveAItis also demonstrated the lowest rate of ambiguous "Need Consensus" outputs (3.4%) compared to experts (16.7%). CONCLUSIONS: UveAItis demonstrated expert-level performance, significantly outperforming general-purpose LLMs and nonexpert human reviewers. These findings validate the potential of domain-specific fine-tuning to enhance the efficiency, scalability, and reproducibility of evidence synthesis in specialized medical fields like ophthalmology. FINANCIAL DISCLOSURES: Proprietary or commercial disclosure may be found in the Footnotes and Disclosures at the end of this article.

Authors

Keywords

No keywords available for this article.