Predicting phenotypes with one step genetic decision trees
Journal:
bioRxiv
Published Date:
Jul 19, 2026
Abstract
Genomic prediction of complex traits is limited when phenotype records are restricted and when using linear models. Increasing the amount of phenotypic data with high-throughput, image-based phenotyping could result in better genomic prediction and stronger signals in variant detection. Here, we analysed phenotypic and genomic data from a selectively bred cohort of the Australasian snapper (Chrysophrys auratus) to identify genetic variants associated with growth traits. We used a high-throughput phenotyping pipeline to extract 13 measurements of size from images. Phenotypic correlations among image-derived and manually measured traits (weight, fork length), together with heritabilities, were analysed. All measurements were significantly positively correlated with each other, and heritability ranged from 0.20-0.38. Genome-wide association studies (GWAS) identified 28 growth-associated SNPs, while GBLUP was used to predict phenotypes, and XGBoost machine-learning models were used to jointly predict phenotypes and report important variants. Both GBLUP (mean R2 = 0.50) and XGBoost (mean R2 = 0.77) performed well on the training data, but performance dropped on testing sets (both = 0.11), which decreased further when accounting for genetic relatedness (both = 0.06). Despite this, approximately 20% of the genetic variance for growth traits was captured by the models, and feature importance from XGBoost reflected signals seen in GWAS. Our findings highlight the utility of integrating computer vision-based phenotyping with GWAS, GBLUP, and ML for trait prediction. Despite detecting shared biological signals as GWAS, genomic prediction faces challenges with population structure and relatedness that are inherent in breeding programmes of mass spawning species, including many aquatic species.