Integrating heterogeneity into topologically associating domain boundary prediction in large genomic context in human.
Journal:
NAR genomics and bioinformatics
Published Date:
Aug 14, 2026
Abstract
A fundamental understanding of genome organization relies on accurately annotating topologically associating domains (TADs) and their boundaries. This is crucial for understanding how cis-regulatory elements regulate gene expression. To go beyond calling TADs and boundaries from Hi-C data, several machine learning-based methods have been proposed to go the step further and predict TAD boundaries from genomic sequences. As the growing evidence of TADs and their boundaries, TADs have been proved exhibiting diverse properties, such as differences in replication timing and epigenetic patterns. However, existing methods do not take this heterogeneity into account. To address this, we propose a method called TADBpred for TAD boundary prediction in a large genomic context in humans. TADBpred focuses on TAD boundaries in active and inactive chromatin across cell-lines and tissues, which are GC-rich and AT-rich, respectively. By integrating genomic elements and sequence composition, we designed two models for GC-rich and AT-rich boundaries, respectively. When testing the performance on respective independent held-out datasets, we obtain AUC scores of 0.91 and 0.80. Our results indicate that TADBpred excels in TAD boundary prediction. Additionally, feature importance analysis highlights the essential features for different classes of TAD boundaries, thereby enhancing our understanding of these TAD boundaries.
Authors
Keywords
No keywords available for this article.