From Public Archive to Reusable Resource: Characterizing Gut Microbiome Metadata in the NCBI SRA
Journal:
bioRxiv
Published Date:
Sep 1, 2026
Abstract
Public sequencing repositories contain large amounts of gut microbiome data that could support cross-study comparison, reproducibility analysis, and microbiome foundation model development. However, the extent to which these data are structured, harmonized, and reusable at archive scale remains unclear. Here, we characterized publicly available gut microbiome sequencing metadata from the NCBI Sequence Read Archive using Google BigQuery, focusing on human gut metagenome, mouse gut metagenome, and broadly annotated gut metagenome records. We evaluated temporal growth, sequencing depth, BioSample and BioProject structure, platform and instrument use, metadata completeness, host attribution, publication linkage, and research themes from linked literature. Public gut microbiome data increased substantially over time and were dominated by human-associated datasets and Illumina sequencing platforms. Core technical metadata fields were highly complete, but biological context needed for reuse, including host identity, phenotype, study design, and disease status, was often inconsistently encoded or required recovery from BioSample attributes and linked publications. In the generic "gut metagenome" cohort, host identity could be assigned for only 13.00% of BioSamples, highlighting the limitations of broad organism annotations for automated cohort construction. Publication linkage was also incomplete at the archive level, although usable text was recovered for most linked publications. Topic modeling of SRA-linked literature showed persistent emphasis on core gut microbiota composition and increasing representation of human cohort and infant microbiome studies. Overall, these findings show that public gut microbiome data are extensive and technically rich but not uniformly analysis ready. Improved metadata harmonization, publication linkage, and biological context recovery will be necessary to support reliable large-scale reuse and AI-ready microbiome data resources.