TENTACLES: a consensus machine learning tool for robust biomarker discovery in heterogeneous data.
Journal:
BioData mining
Published Date:
Jun 10, 2026
Abstract
BACKGROUND: Transcriptomic biomarker discovery often fails to produce reproducible gene signatures across independent cohorts due to model-specific biases and dataset heterogeneity. While single-algorithm approaches may perform well on training data, they frequently fail to generalize effectively. Ensemble methods have proven effective in general machine learning applications, yet their systematic integration for consensus-based feature prioritization remains underexplored in transcriptomics. RESULTS: We developed TENTACLES (Transcriptomic Exploration Tool through Aggregation of Classifiers), an open-source modular framework for robust biomarker discovery through multi-algorithm consensus. The tool is an open-source R package that integrates up to 15 supervised learning algorithms and 6 unsupervised clustering methods. The tool utilizes a modular architecture to automate data preprocessing, multi-algorithm feature prioritization, and cross-cohort validation. By aggregating variable importance across multiple models, TENTACLES identifies gene signatures resilient to algorithm-specific biases. We validated the framework using Crohn's disease as a high-heterogeneity case study across 689 samples from four independent publicly available RNA-seq cohorts. TENTACLES identified a 28-gene consensus panel that achieved superior cross-cohort generalizability compared to single-algorithm-derived signatures and conventional differential expression methods while using, compared to the latter, 95% fewer features. This signature was further refined to a minimal 5-gene core that maintained robust discriminatory power in completely unsupervised validation. These results confirm the tool's ability to extract stable biological signals from complex, noisy datasets. CONCLUSIONS: TENTACLES provides a scalable, disease-agnostic solution for identifying minimal reproducible gene signatures from heterogeneous transcriptomic data. By bridging the gap between complex ensemble modeling and practical biomarker discovery, the software could serve as a versatile resource for researchers aiming to derive reproducible biomarkers across diverse disease contexts.
Authors
Keywords
No keywords available for this article.