Validation is not enough: Longitudinal evidence of post-deployment fragility in clinical AI systems.
Journal:
PLOS digital health
Published Date:
Jul 27, 2026
Abstract
Pre-deployment validation is commonly used to establish the safety and effectiveness of clinical artificial intelligence systems, but acceptable validation performance does not guarantee stable behavior after deployment into routine clinical workflows. We conducted a longitudinal retrospective observational study of four clinically deployed AI systems operating across distinct clinical domains and workflows within a large healthcare organization. Using routinely collected clinical data, outcome labels, and operational telemetry, we compared validation-era performance with post-deployment behavior over extended observation periods. Analyses focused on temporal patterns of discrimination, calibration, data availability, latency, and workflow-related signals, with particular attention to label-dependent and label-independent monitoring. Across all systems, validation-era performance did not persist as a stable operational property after deployment. Calibration drift emerged consistently and often preceded detectable changes in discrimination. Workflow-associated changes in data availability and timing were more strongly and consistently associated with degradation than population-level indicators. Label-independent operational signals, including input missingness and data latency, provided early indication of emerging fragility, whereas outcome-based monitoring was delayed by label latency and documentation processes. These findings suggest that post-deployment fragility can be a structural property of clinical AI systems embedded in evolving workflows. Effective governance therefore requires lifecycle-oriented monitoring strategies that combine calibration reassessment with operational telemetry throughout deployment.
Authors
Keywords
No keywords available for this article.