Synthetic dataset of PLGA and liposome nanocarrier formulations for brain-cancer-relevant drug delivery, release, blood-brain-barrier transport, and paired cell-viability proxies.
Journal:
Data in brief
Published Date:
Jun 1, 2026
Abstract
This data article describes an original synthetic/simulated dataset designed to support materials-informatics and comparative formulation analysis of PLGA nanoparticles and liposomes for brain-cancer-relevant drug delivery. Open row-level datasets that jointly cover formulation descriptors, release kinetics, blood-brain-barrier transport proxies, and paired tumor/non-tumor cell-assay outcomes for PLGA and liposome carriers were not identified by us in a single harmonized open resource during preparation of this package, motivating a transparent synthetic benchmark for methodological and machine-learning reuse. The dataset was inspired by the scientific themes synthesized in the related review article by Makalew and Abrori, but it does not reproduce bibliometric records, published tables, or experimental rows. The package contains 6000 unique virtual formulations in a formulation master table and three linked long-format data tables describing time-resolved release profiles (360,000 rows), blood-brain-barrier-related transport proxies (54,000 rows), and paired tumor/non-tumor cell-assay proxies (432,000 rows), totaling approximately 846,000 assay-like rows. Variables include composition descriptors, preparation routes, physicochemical properties, targeting features, encapsulation efficiency, drug loading, stability, biodegradation proxy, serum stability proxy, integrated blood-brain-barrier transport score, cellular uptake score, biocompatibility score, tumor-directed cytotoxicity proxy, off-target toxicity proxy, and derived multi-criteria performance scores. The synthetic data were generated with a transparent, reproducible workflow that combines domain-informed priors, hierarchical conditional rules, latent heterogeneity, batch effects, replicate variation, bounded noise, sparse scientifically motivated missingness, and post-generation quality filters. The dataset is distributed in open tabular formats together with generation code, a codebook, validation documentation, and reproducible figure-generation scripts.
Authors
Keywords
No keywords available for this article.