Use of machine learning to detect Escherichia coli in drinking water in Bangladesh.
Journal:
PloS one
Published Date:
Aug 20, 2026
Abstract
Escherichia coli (E. coli) is a key indicator of fecal contamination in freshwater and can signal the presence of other harmful bacteria and viruses. The aim of the study is to evaluate the performance of machine learning (ML) tools to detect E. coli in drinking water in Bangladesh using surveillance data under two scenarios: an imbalanced dataset and a balanced dataset. We utilized data from the 2019 Bangladesh Multiple Indicator Cluster Survey, which included a total of 6,069 household drinking water samples. We used agglomerative hierarchical clustering with Ward's linkage to identify district-level hotspots. Extreme Gradient Boosting with SHapley Additive exPlanations values were used for feature selection, and the Synthetic Minority Over-sampling Technique (SMOTE) was used to address class imbalance in the classification task. We applied nine classical ML models in this study: Adaptive Boosting (AdaBoost), Decision Trees (DT), Gradient Boosting Algorithm (GBA), k-Nearest Neighbors (KNN), Light Gradient-Boosting Machine (LightGBM), Logistic Regression (LR), Naïve Bayes (NB), Random Forest (RF), and Support Vector Machine (SVM), along with a Deep Learning Multi-Layer Perceptron (DL-MLP) model to predict the risk of E. coli contamination (REcC) in water. Model performance was evaluated using accuracy, precision, recall, F1 score, Cohen Kappa, area under the curve (AUC), and a violin plot. E. coli contamination in drinking water was detected in 39.2% (95% CI: 37.4-41.2) of households. Bandarban district had the highest REcC. After applying SMOTE and 10-fold cross-validation with hyperparameter tuning, model performance was more consistent across algorithms. In terms of model evaluation, AdaBoost slightly outperformed the others with an accuracy of 81.6%, Cohen kappa statistic of 19.4%, precision of 82.2%, recall of 99%, F1-score of 89.8%, and an AUC of 68.6%. Ensemble model for example AdaBoost and GBA models had ability to accurately classify drinking water samples with respect to the presence of E. coli using surveillance data than others selected model in this study.
Authors
Keywords
No keywords available for this article.