Application of machine learning for identification of key exposure predictors for heavy metal accumulation in hair of traffic police officers in Tehran.
Journal:
The Science of the total environment
Published Date:
Dec 22, 2025
Abstract
In order to determine variability and measure the major exposure factors affecting the levels of hazardous metals (such as Fe, Mn, Ni, Pb, As, Cr, and Cu) in the scalp hair of Tehran traffic police personnel, an advanced statistical method is used. The five most 5 important features have been selected by the Random Forest feature selection technique in order to train the predictive models. In order to identify which features most strongly predict levels of particular metals, we used a variety of supervised machine learning algorithms, such as Linear Regression, Random Forest, and XGBoost. Linear Regression, in conjunction with survey-based demographic, occupational, and lifestyle data and measured metal concentrations. Among the modeling approaches, Linear Regression showed the weakest predictive ability. Its performance was marked by low or even negative R2 values and relatively high RMSE scores across all metals. The metals As (R2 = 0.12), Al (R2 = 0.09), and Fe (R2 = 0.27) calculated low value of R2. Whereas, Random Forest performed considerably better, capturing more variance in metal concentrations, especially for Mn (R2 = 0.77), Fe (0.75), and As (0.73). But, XGBoost outperformed both LR and RF, achieving the highest R2 values for most metals, Pb (0.96), Mn (0.88), Zn (0.83), and Cu (0.81), with remarkably low RMSE values. The comparative results underscore that machine learning techniques, particularly XGBoost, are highly effective in modeling multifactorial data like heavy metal exposure. Our results provide a data-driven foundation for focused occupational health treatments by indicating that variables including age, service duration, mask use, workplace location, and dietary practices (fruit/fish consumption) have varying significance among different metals.