Unsupervised Clustering Enables Cohort-Specific Variant Recalibration in Crops

Journal: bioRxiv
Published Date:

Abstract

Genomic selection (GS) and trait discovery in crops require reliable marker datasets, yet variant discovery remains constrained by the lack of standardised, high-confidence reference resources. Publicly available plant genomic resources are often based on outdated reference genomes, heterogenous pipelines, and generic variant hard-filtering thresholds derived from human studies. Consequently, implementation of the Genome Analysis Toolkit (GATK) Best Practices framework, including machine learning-based variant quality score recalibration (VQSR), remains challenging in crop genomes. Here, we present a crop-adapted framework for unsupervised variant quality assessment and VQSR enabling advanced SNP and insertion-deletion (INDEL) discovery without predefined high-confidence reference variant sets. Using Brassica napus as a model, we use the Illumina 60k Brassica SNP array to generate a reference-specific truth set for data pre-processing and independent evaluation of an unsupervised machine-learning approach for constructing cohort-specific VQSR resources. Gaussian mixture model (GMM)-based clustering of multidimensional variant annotation profiles identifies high-, intermediate-, and low-quality variant classes, providing flexible training resources for VQSR without requiring labelled training data. We further validate the framework in diploid Brassica oleracea and hexaploid wheat (Triticum aestivum) across cohorts differing in population size, genetic diversity, and genome architecture. Despite population- and species-specific differences in variant annotation profiles, GMM-based clustering consistently identifies informative variant quality classes and enables cohort-specific recalibration of both SNPs and INDELs. Our results demonstrate that unsupervised machine learning can extend GATK Best Practices to crop species lacking comprehensive genomic resources, providing a scalable approach for generating high-quality variant datasets for crop genomics and downstream breeding applications.

Authors

  • Bergmann
  • T.; MacNish
  • T. R.; Batley
  • J.; Edwards
  • D.

Categories