Evaluation of an AI-Based Constraint-Optimization Scheduler to Optimize On-Call Schedule Equity and Reduce Administrative Burden in a Pediatric Residency: Retrospective Comparative Study.
Journal:
Journal of medical Internet research
Published Date:
Jul 31, 2026
Abstract
BACKGROUND: Resident scheduling is a high-dimensional optimization problem with implications for workload, fatigue risk, and equity. Real-world evaluations of AI-based constraint-optimization in health care are limited. OBJECTIVE: This study aimed to evaluate an AI-based constraint-optimization scheduler versus a legacy rule-based scheduler for pediatric residency night calls. METHODS: This is a single-center retrospective before-after study at a 235-bed tertiary pediatric center. Twenty-four consecutive months of night-call rosters were analyzed: preimplementation (January to December 2024, legacy rule-based autoscheduler) and postimplementation (January to December 2025, AI-based constraint-programming scheduler combining a local-search metaheuristic solver with human-in-the-loop review). The analytic unit was the resident-month. Outcomes were workload distribution, threshold exceedances (>6 total and >2 weekend calls/month), undesirable sequences (consecutive weekend calls; call-rest-call; call-rest-call-rest-call), equity (mean absolute error from equal share [MAE-ES], root mean square error from equal share), publication lead time, as well as pre- and postsurvey experience. RESULTS: We analyzed 6519 shifts across 1530 resident-months (legacy: 803 resident-months/107 physicians; AI: 727/87; weekend share 28.8% [934/3246] vs 28.6% [935/3273]; service mix P=.99). Mean calls/resident-month did not decline (4.04 vs 4.50; P<.001), but within-period SD was approximately halved. Threshold exceedances fell from 133/803 (16.6%) to 28/727 (3.9%) for >6 calls per month (risk ratio [RR] 0.24, 95% CI 0.15-0.34) and 89/803 (11.1%) to 21/727 (2.9%) for >2 weekend calls (RR 0.27, 95% CI 0.16-0.40; both P<.001). Undesirable sequences declined: consecutive weekends 24.4→18.7/100 resident-months (RR 0.77; P=.02); call-rest-call 51.2→23.4 (RR 0.46); call-rest-call-rest-call 5.6→1.0 (RR 0.18; both P<.001). Equity improved overall and within every qualification stratum: MAE-ES -0.26 shifts (95% CI -0.28 to -0.23) and RMSE-ES -0.29 (95% CI -0.32 to -0.26); Senior, Advanced, and Novice strata were all P<.001 after Holm correction. Publication lead time more than doubled (10.7→21.2 d; Δ+10.5, Cohen d=4.78; Cliff δ=1.00; P<.001). Interrupted time-series confirmed immediate level shifts for >6-call exceedances (β=-8.88; P=.004), MAE-ES (β=-0.18; P<.001), call-rest-call (β=-13.17; P=.002), call-rest-call-rest-call (β=-2.35; P=.03), and >2 weekend exceedances (β=-7.32; P<.001), with stable postimplementation fairness slopes. Among survey respondents (n=47 pre; n=38 post), software satisfaction rose 6.77→8.71/10, perceived timeliness 3.28→4.61/5, perceived consecutive-night frequency 3.15→4.24, and perceived equity 2.98→3.61 (all P≤.006). CONCLUSIONS: An AI-based constraint-optimization scheduler was associated with significantly more equitable on-call workload across all qualification strata, large reductions in high-risk shift sequences and threshold exceedances, and a doubling of publication lead time, despite no reduction in mean per-physician burden once all physicians were retained. Multisite prospective replication is warranted before generalization.
Authors
Keywords
No keywords available for this article.