Deploying Local Large Language Models for Automated Article Screening in Scientific Literature Reviews
Journal:
medRxiv
Published Date:
Oct 6, 2026
Abstract
Background Large language models (LLMs) show promise in scientific literature reviews, but cloud-based proprietary models like GPT-4 raise concerns around privacy, copyright, and licensed article access. Using a scoping review on AI applications for treatment effect estimation as a use case, this study evaluated local open-source LLMs for automated article screening in scientific literature reviews and developed a practical workflow for their implementation. Methods We extracted a sample dataset of 300 records from an ongoing scoping review on machine learning methods for treatment effect estimation using real-world medical data. We divided the dataset into training (n = 200) and testing (n = 100) datasets. We evaluated 11 locally deployed open-source LLMs from five model families (GPT-OSS, Gemma, Qwen, Mistral, Llama) for both abstract screening and full-text screening. Full-text screening was done using text converted by two PDF-to-text conversion tools: Docling and PyMuPDF4LLM. Performance was assessed using raw accuracy with 95% Wilson's confidence interval, relaxed accuracy, Cohen's kappa, and runtime. During workflow development, prompt engineering, output extraction strategies, and generation settings were iteratively optimized using the training dataset. Results GPT-OSS-20B achieved the highest raw accuracy for title and abstract screening (0.890 on the training set and 0.850 on the test set), followed closely by Gemma-4-31B-it (0.875 on the training set and 0.830 on the test set) and Gemma-4-26B-A4B-it (0.810 on the training set and 0.790 on the test set). In full-text screening, slightly higher raw accuracy and reduced false-positive rates were observed, with GPT-OSS-20B and Gemma-4 variants consistently demonstrating strong performance. Only minor variations were found between the two PDF-to-text conversion methods evaluated, though runtime varied substantially across models. Based on these findings, we compiled a set of practical recommendations for local LLM-assisted article screening. Conclusions Local LLMs can effectively support literature screening when integrated into a carefully designed workflow. Recent generations of open-source local models achieved performance comparable to human reviewers and proprietary LLMs. Local deployment also provides advantages in data privacy, copyright compliance, customization, and cost-efficiency. While human oversight remains essential, the proposed workflow demonstrated that local LLMs represent a practical and reproducible approach for accelerating evidence synthesis.