Machine learning prediction of soil adsorption coefficients for industry-related aromatic contaminants: Molecular drivers and global distribution insights.
Journal:
Journal of hazardous materials
Published Date:
Jun 2, 2026
Abstract
Predicting the soil adsorption behavior of industry-related aromatic contaminants (ACs) is crucial for assessing their environmental fate and risks. However, existing models that rely on limited parameters, such as hydrophobicity, are often insufficient, particularly for emerging ACs, as their structural diversity introduces complex adsorption mechanisms (e.g., electrostatic interactions and polar effects) that cannot be adequately captured by static hydrophobicity descriptors. In this study, a global dataset of 3044 data points from 483 adsorption isotherms, covering 114 ACs in 311 soils, was compiled. Machine learning (ML) models were developed using soil properties and four types of molecular representations to predict the soil adsorption coefficient (Kd). Among them, the Support Vector Machine model using RDKit descriptors achieved the best performance in cross-validation (R2 = 0.84, MAE = 0.37, RMSE = 0.50) and testing set (R2 = 0.89, MAE = 0.33, RMSE = 0.44). SHapley Additive exPlanations analysis identified, in addition to the equilibrium concentration and soil properties, molecular descriptors related to polarity, surface charge distribution, and drug-likeness strongly influencing Kd. Applicability domain analysis showed that 71% of 552 environmentally detected ACs fell within the reliable prediction space. Furthermore, global distribution of soil adsorption capacities for representative ACs was predicted. Emerging ACs (pentachlorobenzene and tetrabromobisphenol A) exhibited generally higher soil adsorption capacities than traditional ACs (naphthalene and phenanthrene). These global distributions of Kd values and key molecular drivers provide valuable insights into the environmental behaviors and risk management of ACs.
Authors
Keywords
No keywords available for this article.