Machine Learning-Driven Phenotype Predictions based on Genome Annotations
Journal:
bioRxiv
Published Date:
Oct 7, 2026
Abstract
The rapid expansion of isolate genome sequences and metagenome-assembled genomes has created a growing need for computational approaches that can infer microbial phenotypes directly from genome information. Here, we present a machine learning framework that uses RAST-derived functional annotations as predictive features to classify bacterial phenotypes. We evaluated six machine learning algorithms; K-Nearest Neighbors, Gaussian Naive Bayes, Support Vector Machines, Neural Networks, Logistic Regression, and Decision Trees across three phenotype-prediction tasks: Gram-stain type (Gram-negative and Gram-positive), respiration type (aerobic, anaerobic, and facultative anaerobic), and growth phenotype (growth and no growth). All six classifiers showed strong predictive performance, although performance varied by phenotype. Gram-stain type was predicted with near-perfect performance, with precision and recall exceeding 0.99 across classifiers. Respiration type was predicted with a maximum overall accuracy of 81.5% using Gaussian Naive Bayes, while growth versus no growth on individual carbon substrates reached approximately 86% accuracy. Decision-tree models additionally provided interpretable links between phenotype predictions and specific functional roles, enabling biological interpretation of the features underlying classification. To make the approach broadly accessible and reproducible, we implemented a four-application workflow in the U.S. Department of Energy Systems Biology Knowledgebase (KBase) that enables users to: (i) upload high-quality data to train classifiers; (ii) annotate genomes in the training set using the RAST annotation algorithm; (iii) build six different genome classifiers; and (iv) predict the phenotypes of unclassified genomes. Together, these results demonstrate that genome annotations provide an effective and interpretable feature space for machine learning-based microbial phenotype prediction and establish a reusable framework for extending this approach to additional microbial traits.