Geometric vs. global knowledge in GNNs: a controlled leakage study for early-late TS classification of the isomerization of the 2-norbornyl carbocation.

Journal: Physical chemistry chemical physics : PCCP
Published Date:

Abstract

Transition-state classification remains challenging because distinguishing reactant-like from product-like geometries requires resolving subtle mechanistic changes. Although GNNs can learn molecular representations, it remains unclear whether their predictions arise from genuine mechanistic information in geometry or from global descriptors that inadvertently encode reaction-coordinate information. A controlled leakage study was performed on the highly interconnected 2-norbornyl carbocation isomerization network, which contains overlapping pathways and complex rearrangement features. A hierarchical series of GNNs was constructed to progressively introduce increasingly informative global descriptors. These descriptors ranged from displacement and force-based metrics to structural-distance relationships and normalized reaction-coordinate fractions. The geometry-only GNN performed near randomly, indicating that local graph representations alone are insufficient for robust early-late TS classification. Adding displacement-based descriptors substantially improved performance, while force-based and redundant descriptors provided only modest additional gains. A sharp increase to near-perfect accuracy occurred after introducing reactant-TS and product-TS structural-distance descriptors, even before explicitly including the normalized reaction fraction. Multiple analyses showed that these chemically meaningful global descriptors can unintentionally encode target-defining reaction-progress information, highlighting the need for systematic leakage detection in graph-based chemical machine learning. The inclusion of electronic properties, such as Mayer bond orders and Mulliken charges increased model accuracy without data leakage. More broadly, this work establishes a systematic framework for identifying and quantifying information leakage in graph-based chemical machine learning.

Authors

Keywords

No keywords available for this article.