Benchmarking computational methods for multi-omics biomarker discovery in cancer

Journal: bioRxiv
Published Date:

Abstract

Multi-omics profiling characterizes cancer biology and supports biomarker discovery for prognosis and therapy selection. Although numerous computational methods have been proposed, their ability to identify clinically relevant biomarkers has not been systematically evaluated. Here, we benchmark 20 representative statistical, machine learning (ML) and deep learning (DL) methods using curated gold-standard prognostic and therapeutic biomarkers across five real-world datasets. We evaluate performance in terms of both biomarker identification accuracy and stability. Pathway-informed transformer methods, including DeePathNet and DeepKEGG, achieve the best overall performance. Across methods, effective biomarker recovery is associated with the integration of biological knowledge, global feature interactions, multivariate feature attribution, and effective regularization. Analysis of omics-type contributions reveals method- and modality-specific biases, highlighting the importance of broader omics integration. We further evaluate methods on simulated datasets to probe sensitivity with controlled signal and noise. Finally, by aggregating results from top-performing methods, we construct consensus biomarker panels that nominate candidates for future investigations. All datasets, code, and evaluation pipelines are released as a public resource to support future studies in multi-omics biomarker discovery for cancer.

Authors

  • Athan Z. Li; Yuxuan Du; Yan Liu; Liang Chen; Ruishan Liu