VN1K is a pangenome-informed multi-omics and phenomics resource for the Vietnamese population.
Journal:
Nature communications
Published Date:
Jul 22, 2026
Abstract
The population of Vietnam remains underrepresented in global genomic databases. Here, we present VN1K, a resource of multi-omics and phenotypic information for 1011 unrelated Vietnamese individuals. We present high-depth short-read whole-genome sequencing data for all samples along with various -omics datasets. Using a high-sensitivity variant detection pipeline, which includes a pangenome graph reference and a deep-learning framework, we identify approximately 42 million variants with 7 million short insertions/deletions and 90 thousand structural variants. VN1K also features a whole-genome methylation profile based on long read sequencing. We create a genotype imputation panel with high accuracy on the Vietnamese population, allowing us to identify variants with significantly different allele frequencies in the Vietnamese population compared to other populations. We establish the functional relevance of some of these variants, particularly those in genes associated with genetic disorders, immune diseases, and drug responses, by integrating the allele frequency differences with known genotype-phenotype associations and clinical annotations. Further, we map various loci related to hepatitis B virus infection, triglyceride levels, LDL-C levels, serum glucose levels, HbA1c levels, and levels of two liver enzymes (ALT and AST). The VN1K dataset is accessible via genome.vinbigdata.org, an integrated platform with both linear and graph-based genome browsers.
Authors
Keywords
No keywords available for this article.