Multi-variant GWAS and genomic prediction dissect the genetic architecture underlying tomato fruit weight.

Journal: TAG. Theoretical and applied genetics. Theoretische und angewandte Genetik
Published Date:

Abstract

Tomato fruit weight is primarily controlled by stable genetic effects and domestication-/improvement-associated loci, including prominent chromosome 5 signals, and integrating SNP, INDEL, and SV diversity improves candidate-locus discovery, biological interpretation, and genomic prediction across environments. Tomato fruit weight (fw) is a major breeding target shaped by domestication, crop improvement, and environmental variation. We investigated the genetic architecture and genomic predictability of fw in the Varitome population representing the tomato domestication continuum, including Solanum pimpinellifolium (SP), S. lycopersicum var. cerasiforme (SLC), and cultivated S. lycopersicum (SLL). Genome-wide SNPs, insertions and deletions (INDELs), and structural variants (SVs) were analyzed individually and in combination to assess their contributions to association mapping, candidate locus discovery, and genomic prediction. Population structure based on all three variant classes clearly separated the domestication groups. Genome-wide association analyses identified 15 SNP, 10 INDEL, and 10 SV loci significantly associated with fw across environments. Several associations co-localized with known fw genes, including fw2.2/CNR, fw3.2/KLUH, fw11.3/CSR, lc/WUSCHEL, and fas/CLAVA3 supporting the biological relevance of the detected signals. Chromosome (Chr) 5 was particularly notable, with a 390-bp deletion at chr5_45,551,023 detected across all four environments and an INDEL at chr5_55,361,233 detected across three environments, suggesting stable improvement-associated candidate loci for fw variation validation. Multiple loci exhibited clear allele-frequency shifts from SP through SLC to SLL, consistent with selection during domestication and improvement, whereas others were detected almost exclusively in cultivated germplasm, suggesting more recent breeding-associated origins. Fw displayed strong genetic control, with genotype explaining more than 92% of the phenotypic variance and only minor contributions from environment and genotype-by-environment interactions. Five genomic prediction models based on SNPs, INDELs, SVs, their combinations, and genotype-by-environment interaction kernels were evaluated using four cross-validation scenarios (CV1, CV2, CV0, and CV00). Prediction accuracy was influenced more strongly by validation scenario and prediction algorithm choice than by marker class. Partial Least Squares (PLS) and Bayesian Genomic Linear Regression (BGLR) consistently outperformed Random Forest (RF) and Deep Learning (DL) models, particularly under scenarios involving untested genotypes and environments. INDEL-based models frequently achieved the highest accuracies under the most stringent prediction scenarios, although differences among marker classes were generally modest. Integrating multiple classes of genomic variation improved biological interpretation, enabled the identification of domestication- and improvement-associated loci, and provided accurate prediction across environments and genetic backgrounds. Together, these results demonstrate that tomato fw is governed primarily by stable additive genetic effects and highlight the value of multi-variant genomic approaches for both genetic dissection and genomic-assisted breeding.

Authors

Keywords

No keywords available for this article.