Machine learning-based classification model to distinguish tumor and tumor-free samples using synthetic plasma proteomic dataset.

Journal: La Clinica terapeutica
Published Date:

Abstract

BACKGROUND: Tumor-Specific Peptides (TSPs) and Tumor-associated Overexpressed Proteins (TOPs) are promising biomarkers for cancer diagnosis and monitoring. TSPs arise from somatic mutations unique to tumor cells. Their detection in plasma through mass spectrometry-based proteomics provides a non-invasive alternative to traditional biopsy-based diagnostics. We developed a machine learning workflow to classify tumor versus tumor-free samples using synthetic proteomic features. METHODS: A synthetic (in silico) dataset was generated, parameterized using curated proteogenomic resources (CAPD, CPTAC, dbPepNeo2.0) and a literature-based search. For model development, 52 TSP detection indicators and 5 TOP concentrations (AFP, CEA, PSA, VEGF, ADFP) were simulated for 400 individuals, equally distributed between tumor-affected and healthy. No real patient-level data or biological samples were used. Multiple supervised classifiers were trained and validated using 5-fold stratified cross-validation (random seed = 42). RESULTS: In cross-validation, Support Vector Machine (SVM) achieved the highest mean recall (0.88 ± 0.10) and AUC (0.962 ± 0.029) and was selected as the Stage-2 classifier. On the holdout test set (n=80), the two-stage TSP gate + Stage-2 ML strategy achieved accuracy 0.8875, recall 0.90, and a composite-score AUC of 0.924 (score=1.0 if any TSP is detected; otherwise, the Stage-2 predicted probability). CONCLUSION: This in silico proof-of-concept suggests that integrating mass spectrometry-inspired peptide detection with machine learning could support future plasma-based screening and recurrence monitoring. Challenges include plasma proteome complexity and the low abundance of tumor-derived peptides; real-world validation is essential for clinical transition.

Authors

Keywords

No keywords available for this article.