Predicting Call Abandonment in a Health Care Call Center Using Nonpersonal Operational Data: Machine Learning Study.
Journal:
JMIR medical informatics
Published Date:
Aug 31, 2026
Abstract
BACKGROUND: Call abandonment is a critical barrier to patient access in health care call centers; yet, predictive modeling efforts are limited by strict privacy regulations that restrict the use of personal or behavioral data. Whether abandonment can be accurately predicted using only anonymized operational metrics remains unclear. OBJECTIVE: This study evaluated the feasibility, performance, and operational use of machine learning models trained exclusively on nonpersonal, routinely collected call center metrics to predict call abandonment across distinct organizational phases. METHODS: We analyzed 1,037,363 call records from a large academic health care system spanning 4 operational periods marked by workflow changes and skill consolidation. Features included temporal variables, skill identifiers, and rolling operational metrics (in-queue time, occupancy, handle time, after-call work time, and active agents). Random forest and CatBoost models were trained on 3 phases defined as: (T1 [January-April 2023; original workflows], T2 [May-August 2023; post skill consolidation, cross-training, and new workflows], and T3a [September-December 2023; optimized processes]) using 5-fold cross-validation with 3 imbalance-handling strategies (none, class weighting, and synthetic minority oversampling technique). Temporal generalizability was assessed by evaluating all models on all 4 phases, with T3b serving as an unseen holdout set. Performance was evaluated using area under the curve (AUC), precision-recall area under the curve (PR-AUC), Brier score, and calibration error. Shapley additive explanations (SHAP) values quantified feature contributions. RESULTS: Across all training phases and algorithms, adding operational metrics improved the area under the receiver operating characteristic curve by 0.03-0.13 vs models using only temporal and skill features. The best configuration was a CatBoost model trained on T2 with operational metrics and no imbalance correction (T3b AUC 0.767, PR-AUC 0.068, expected calibration error 0.006, and Brier 0.027). Models trained solely on preintervention data (T1) generalized poorly to postintervention periods when restricted to temporal and skill features (AUC 0.32-0.38) but achieved AUC 0.715 on T3b when operational metrics were included. SHAP analysis consistently identified in-queue time as the dominant predictor, with the number of logged-in agents, hour-of-day, and skill identifiers comprising the remaining top features. Abandonment declined from 8.7% (20,809/238,722) in T1 to 2.8% (8293/300,060) in T3b; model-based analyses of temporal features (day of week and hour of day) showed the highest risk on Mondays and between 11:00 and 16:00. Skill-level analyses showed marked improvement in high-volume imaging teams with high abandonment. CONCLUSIONS: Models trained on nonpersonal operational data predicted call abandonment on the T3b temporal holdout, with real-time queue and staffing metrics providing the dominant signal. These findings support the feasibility of interpretable, operationally grounded abandonment prediction in health care call centers and indicate that models should be retrained after major operational changes.
Authors
Keywords
No keywords available for this article.