Pneumonia Detection in Paediatric Chest X-Rays using Ensembled Large Language Models
Journal:
medRxiv
Published Date:
Sep 5, 2026
Abstract
Background: Paediatric pneumonia is a major cause of childhood morbidity and mortality. Chest X-rays (CXR) are central to diagnosis, but shortages of specialist radiologists can delay reporting. Multimodal large language models (MLLMs) may assist clinical workflows by analysing images and communicating findings, but few open-source MLLMs are specifically pre-trained on paediatric CXRs. Objective: To evaluate whether ensemble strategies improve MLLM diagnostic performance for paediatric radiological pneumonia detection on CXRs without any fine-tuning, and to assess the generalisability of ensemble findings across datasets. Methods: In this retrospective study, paediatric CXRs from three datasets were analysed. Two internal datasets (balanced and real-world) were acquired from KK Women's and Children's Hospital (KKH), Singapore, where images were independently reviewed by two board-certified radiologists with pneumonia severity assigned to three classes using a predefined consensus algorithm. A third external dataset, the publicly available Kermany CXR Pneumonia dataset, was included to evaluate generalisability. Fifteen MedGemma-4B-it agents classified each CXR, with majority voting, soft voting, and GPTOSS-20B aggregation compared against baseline average agent performance. The primary outcome was One-vs-Rest (OvR) AUROC. Results: The KKH balanced dataset contained 1,000 CXRs and the real-world dataset contained 1,300 CXRs. The Kermany external validation dataset contained 5,856 CXRs. Soft voting significantly improved OvR-AUROC over baseline in all three datasets: balanced dataset improved by +0.065 (95% Confidence Interval (CI) = [+0.051, +0.078], P = 0.0001), real-world dataset improved by +0.073 (95% CI = [+0.051, +0.091], P = 0.0003), and Kermany dataset improved by +0.103 (95% CI = [+0.087, +0.119], P = 0.0001). In the KKH datasets, soft voting also improved accuracy, Cohen's {kappa}, and One-vs-One (OvO) AUROC, and improved F1-score in the balanced dataset. GPTOSS-20B and majority voting demonstrated significantly superior specificity across all datasets. Conclusion: Soft voting ensembles consistently enhance MedGemma's diagnostic discriminatory performance for paediatric radiological pneumonia detection across both internal and external validation datasets. This ensemble framework enables privacy-preserving, near real-time clinical decision support with explainable outputs, with potential for integration into paediatric emergency workflows pending prospective validation. The system's high specificity supports a triage role in flagging high-risk radiological pneumonia cases for urgent clinician review.