Diagnostic performance of ¹⁸F-FDG PET/CT radiomics-based machine learning models for pancreatic lesion characterisation: Comparison with visual assessment and evaluation of human-radiomics synergy.

Journal: European journal of nuclear medicine and molecular imaging
Published Date:

Abstract

PURPOSE: We aimed to evaluate the diagnostic performance of ¹⁸F-FDG PET/CT-derived radiomics-based machine learning models for differentiating malignant from benign pancreatic lesions, to compare these models with two sequential stages of visual assessment, and to assess whether incorporation of clinician judgement as a model input provides additional diagnostic gain. METHODS: In this retrospective single-centre study, 853 consecutive patients who underwent ¹⁸F-FDG PET/CT between April 2009 and February 2025 for known or suspected pancreatic lesions were screened, and 466 were included in the final analysis. Final diagnosis was established by histopathology (76.9%) or clinico-radiological follow-up (23.1%). Diagnostic performance was assessed within a five-stage framework consisting of two visual assessment stages and three progressively expanded machine-learning stages. The visual stages comprised first-look assessment based on PET/CT images alone and comprehensive second-look assessment integrating available clinical, laboratory, and multimodal imaging data. The machine-learning stages comprised a hybrid PET + CT radiomics model, an integrated model additionally incorporating clinical, laboratory, semiquantitative PET, and imaging-derived variables, and a human-radiomics synergistic model including the clinician's second-look binary judgement as an input feature. PET-only and CT-only radiomics models were also evaluated as unimodal comparators. Radiomic features were extracted after manual three-dimensional segmentation using 3D Slicer/PyRadiomics, and Random Forest was selected as the final classifier. RESULTS: Of 466 lesions, 336 (72.1%) were malignant and 130 (27.9%) were benign. First-look visual assessment achieved 93.5% sensitivity, 49.2% specificity, and 81.1% accuracy. Comprehensive second-look assessment improved performance to 97.9% sensitivity, 76.9% specificity, and 92.0% accuracy. Among radiomics-based models, the hybrid PET + CT model outperformed unimodal approaches, achieving 91.1% sensitivity, 66.7% specificity, and 86.8% accuracy. The integrated model did not improve overall accuracy beyond the hybrid model. The human-radiomics synergistic model achieved the highest sensitivity (98.7%) and overall accuracy (93.3%), whereas comprehensive second-look assessment retained slightly higher specificity and NPV. CONCLUSION: ¹⁸F-FDG PET/CT radiomics-based machine learning models improved specificity beyond routine first-look visual PET interpretation, approaching the performance of comprehensive second-look visual assessment without requiring additional clinical data. Incorporation of clinician judgement further increased sensitivity, although comprehensive second-look visual assessment by an experienced reader retained the highest specificity and NPV among all approaches. These findings support a complementary, decision-support role for radiomics-based models alongside, rather than in place of, expert visual assessment.

Authors

Keywords

No keywords available for this article.