A Detection, Diagnostic, and Triage AI Buyer's Guide: Questions that Matter.

Journal: Journal of imaging informatics in medicine
Published Date:

Abstract

The published performance of artificial intelligence (AI) models in radiology is typically based on the reporting of sensitivity, specificity, and receiver operating characteristic area under the curve, both in the peer-reviewed literature and for Food and Drug Administration 510(k) submissions. Interestingly, these metrics cannot inform radiologists, the users of these systems, of the rate or quantity of each type of error to anticipate if they implement the candidate AI product(s) into their own clinical practice. Only the positive predictive value (PPV) and negative predictive value (NPV), or rather, their complements, the false discovery rate (1-PPV, FDR), and the false omission rate (1-NPV, FOR) can provide these error rates. Although some published articles and 510(k) submissions include the PPV and NPV of the concerned AI models, many do not, and the ones that do sometimes test on enriched datasets with artificially high prevalence rates of the target condition, thus inflating PPV and deflating NPV that would be found in clinical practice. This manuscript demonstrates how clinical practices can estimate and evaluate an AI's FDR and FOR for their clinical population using Bayes' Theorem. We also propose a risk-based evaluation matrix (RADDE) which allows radiologists to consider the medical, legal, financial, workflow, psychological, and reputational impact of these estimated AI error rates.

Authors

Keywords

No keywords available for this article.