OmicsMLMentor: A Web Application for Guided Machine Learning Analysis of Omics Data.
Journal:
Journal of proteome research
Published Date:
Apr 28, 2026
Abstract
Expression-based omics technologies (e.g., proteomics, metabolomics, transcriptomics, etc.) increasingly rely on supervised and unsupervised machine learning (ML) models to find key biomolecules distinguishing conditions, identify natural groupings in biological data, or generate predictions for outcomes of interest. Fitting ML models to omics data presents several challenges, including handling missing data, selecting a normalization method, choosing a valid model, and optimizing hyperparameters, all requiring statistical programming skills to address these challenges. Thus, the open-source web application OmicsMLMentor was designed to lower the barrier to ML modeling for omics data. OmicsMLMentor supports the fitting of 15 ML models (10 supervised and 5 unsupervised) tailored to omics data sets, such as proteomics, metabolomics, lipidomics, and transcriptomics. OmicsMLMentor offers several omics-specific features, including methods for handling missingness (imputation, conversion, removal), normalization tests, ranking of models based on the structure of a user's data and user input, and optimal hyperparameter selections using cross-validation splits. By streamlining ML workflows for omics analysis, OmicsMLMentor addresses critical gaps in existing online web tools, facilitating a broader adoption of these models for omics research. Here, OmicsMLMentor is applied to data from a lignin exposure study to highlight example workflows for fitting both supervised and unsupervised models to data.
Authors
Keywords
No keywords available for this article.