Genome-wide transcriptomics and a machine learning-based random forest model identify novel biomarkers to predict phenotypic variability in Wilson disease.

Journal: Human immunology
Published Date:

Abstract

Genotype-phenotype correlation studies fail to explain the phenotypic diversity in Wilson disease (WD) and the missing link may lie in transcriptomics. Therefore, we performed genome-wide transcriptomic profiling and machine-learning-based random forest modeling (RF) to identify the most important genes that could explain phenotypic variability in WD and reveal the molecular pathogenesis of WD. For this purpose, we used bulk RNA sequencing approaches to include 47 genetically confirmed pediatric hepatic WD patients and controls. We observed 321 differentially expressed genes (DEGs) in WD. Among top 100 DEGs, 42 significant DEGs (p < 0.05) were correlated with Kayser-Fleischer ring, serum ceruloplasmin, 24-h urinary copper levels, PELD score at recruitment, and degree of presentation at diagnosis. The random forest model identified a panel of ten genes: GPR37L1, OIP5, TRAJ23, IGHA2, TDRD1, LILRA4, HMOX2, UGT2B17, IGLV7-46 and EPHA1 that could explain 97.22% of phenotypic variability in WD. Ingenuity pathway analysis ranked MAP kinase pathway highest, driven by upregulation of RPS6KA1 (log2FC = 4.68) and downregulation of DUSP16 (log2FC = -5.22), and GAPLINC (log2FC = -2.45). Iron homeostasis signaling demonstrated comparable enrichment, involving key genes ATP6V0D2 (log2FC = -4.06), HMOX2 (log2FC = -2.70), HBD (log2FC = 1.96) and HBZ (log2FC = 1.76). These results suggest that children with WD have dysregulation in MAP Kinase and iron homeostasis signaling pathways providing an important insight into the molecular pathogenesis of WD. RF modeling identified a panel of biomarker with ten genes, that could explain the phenotypic variability in WD.

Authors

Keywords

No keywords available for this article.