Generation of Molecules Near the Applicability Domain Boundaries of Property Prediction Models.
Journal:
Journal of chemical information and modeling
Published Date:
Jun 4, 2026
Abstract
In the molecular design of high-performance materials, the physical properties (y) of chemical structures generated virtually by a computer is predicted using a machine learning model, and chemical structures that satisfy the target value of y are then selected. Although the potential for improving the y value is higher outside the applicability domain (AD) of the model than inside the AD, the predicted values of y are unreliable, and chemical structures with improved y values near the AD boundary are explored while maintaining the reliability of the prediction. The aim of this study is to efficiently generate chemical structures near the AD boundary. Among the chemical structures generated using a generative adversarial network (GAN), the chemical structures close to the AD boundary were swapped with the chemical structures in the GAN training data that were farthest from the AD boundary, and the GAN was retrained. Validation using three physical property data sets (the boiling point, melting point, and water solubility) confirmed that the proposed method generates a higher proportion of chemical structures near the AD boundary than conventional GANs.
Authors
Keywords
No keywords available for this article.