Uncertainty aware training to improve deep learning model calibration for classification of cardiac MR images.

Journal: Medical image analysis

Published Date: Jun 1, 2023

Abstract

Quantifying uncertainty of predictions has been identified as one way to develop more trustworthy artificial intelligence (AI) models beyond conventional reporting of performance metrics. When considering their role in a clinical decision support setting, AI classification models should ideally avoid confident wrong predictions and maximise the confidence of correct predictions. Models that do this are said to be well calibrated with regard to confidence. However, relatively little attention has been paid to how to improve calibration when training these models, i.e. to make the training strategy uncertainty-aware. In this work we: (i) evaluate three novel uncertainty-aware training strategies with regard to a range of accuracy and calibration performance measures, comparing against two state-of-the-art approaches, (ii) quantify the data (aleatoric) and model (epistemic) uncertainty of all models and (iii) evaluate the impact of using a model calibration measure for model selection in uncertainty-aware training, in contrast to the normal accuracy-based measures. We perform our analysis using two different clinical applications: cardiac resynchronisation therapy (CRT) response prediction and coronary artery disease (CAD) diagnosis from cardiac magnetic resonance (CMR) images. The best-performing model in terms of both classification accuracy and the most common calibration measure, expected calibration error (ECE) was the Confidence Weight method, a novel approach that weights the loss of samples to explicitly penalise confident incorrect predictions. The method reduced the ECE by 17% for CRT response prediction and by 22% for CAD diagnosis when compared to a baseline classifier in which no uncertainty-aware strategy was included. In both applications, as well as reducing the ECE there was a slight increase in accuracy from 69% to 70% and 70% to 72% for CRT response prediction and CAD diagnosis respectively. However, our analysis showed a lack of consistency in terms of optimal models when using different calibration measures. This indicates the need for careful consideration of performance metrics when training and selecting models for complex high risk applications in healthcare.

Authors

Tareen Dawood

School of Biomedical Engineering & Imaging Sciences, King's College London, UK. Electronic address: tareen.dawood@kcl.ac.uk.
Chen Chen

The George Institute for Global Health, Faculty of Medicine, University of New South Wales, Sydney, NSW, Australia.
Baldeep S Sidhu

School of Biomedical Engineering & Imaging Sciences, King's College London, UK; Guy's and St Thomas' Hospital, London, UK.
Bram Ruijsink
Justin Gould

School of Biomedical Engineering & Imaging Sciences, King's College London, UK; Guy's and St Thomas' Hospital, London, UK.
Bradley Porter

School of Biomedical Engineering & Imaging Sciences, King's College London, UK; Guy's and St Thomas' Hospital, London, UK.
Mark K Elliott

School of Biomedical Engineering & Imaging Sciences, King's College London, UK; Guy's and St Thomas' Hospital, London, UK.
Vishal Mehta

School of Biomedical Engineering & Imaging Sciences, King's College London, UK; Guy's and St Thomas' Hospital, London, UK.
Christopher A Rinaldi

Division of Imaging Sciences and Biomedical Engineering, King's College London, London, United Kingdom.
Esther Puyol-Anton
Reza Razavi
Andrew P King

Division of Imaging Sciences and Biomedical Engineering, King's College London, London, United Kingdom. Electronic address: andrew.king@kcl.ac.uk.

Keywords

Artificial Intelligence Calibration Coronary Artery Disease Deep Learning Heart Humans Uncertainty

External Resources

View on PubMed Access via DOI PubMed (37327613)

Uncertainty aware training to improve deep learning model calibration for classification of cardiac MR images.

Abstract

Authors

Keywords

External Resources

Popular Topics

Recent Journals