Can machine learning uncover viable non-platinum catalysts for hydrogen evolution? An unsupervised clustering approach to transition metal-ligand complexes.
Journal:
RSC advances
Published Date:
Jul 22, 2026
Abstract
Hydrogen is central to the global clean energy transition, yet the widespread use of green hydrogen is constrained by its dependence on costly, supply-limited platinum-based catalysts in Proton Exchange Membrane (PEM) electrolyzers. While alternatives have been explored, identifying efficient and stable non-platinum catalysts remains a major challenge due to the vast chemical space and the expense of experimental screening. This study pioneers a novel unsupervised machine learning framework to systematically map this vast chemical space and identify promising catalyst candidates without reliance on pre-existing experimental data. Using the tmQM dataset comprising over 108 000 transition metal-ligand complexes; we selected five key descriptors (HOMO-LUMO gap, dipole moment, molecular size, metal node degree, and total charge) to represent catalytic potential and stability. Three unsupervised clustering algorithms (K-Means, DBSCAN, and Gaussian Mixture Models) were applied to identify chemically and electronically distinct groups. Unlike supervised models that predict properties for known structures, our unsupervised approach uncovers previously unidentified classes of non-platinum complexes that possess features ideal for catalytic activity. This work establishes a new, efficient workflow for catalyst science, shifting the focus from slow, one-by-one candidate evaluation to a rapid, holistic mapping of chemical space to pinpoint the most promising regions for focused experimental investigation. Importantly, all shortlisted complexes are structurally documented in crystallographic databases, which enables direct experimental validation and minimizes trial-and-error in material design.
Authors
Keywords
No keywords available for this article.