Leveraging Chemical Hidden-Space Representations Effectively in Bayesian Optimization for Experiment Design through Dimension-Aware Hyperpriors.

Journal: Journal of chemical theory and computation
Published Date:

Abstract

Hidden-space molecular representations derived from pretrained graph neural networks, transformer models, and large language models offer an appealing alternative to conventional encodings for Bayesian optimization (BO) in chemical experiment design. However, their performance benefits in BO remain poorly understood and difficult to assess in practice, and when deployed within current chemical BO workflows, they often do not yield clear advantages over traditional representations. Here, we show that this behavior arises from a previously underappreciated interaction between representation dimensionality and Gaussian process kernel hyperpriors, rather than from limitations of the representations themselves. We demonstrate that different chemical representations induce substantial variations in search-space dimensionality, which, when paired with fixed or mismatched length scale hyperpriors, lead to flat marginal likelihood landscapes and severely degrade surrogate learning and acquisition optimization. To address this issue, we introduce a dimension-aware adaptive hyperprior whose characteristic length scale scales with the square root of the search-space dimensionality. Benchmarking across multiple reaction optimization data sets and molecular representations shows that this adaptive prior consistently restores and amplifies the performance advantages of hidden-space featurizations over traditional representations, such as Mordred or one-hot encoding. Our results identify hyperprior calibration as a critical requirement for the fair evaluation and effective deployment of modern molecular representations in chemical BO.

Authors

Keywords

No keywords available for this article.