Comparative Evaluation of Deep Learning and Foundation Model Embeddings for Osteoarthritis Feature Classification in Knee Radiographs.

Journal: Journal of imaging informatics in medicine
Published Date:

Abstract

Foundation models (FM) offer a promising alternative to supervised deep learning (DL) by enabling greater flexibility and generalizability without relying on large, labeled datasets. This study investigates the performance of supervised DL models and pre-trained FM embeddings in classifying radiographic features related to knee osteoarthritis. We analyzed 44,985 knee radiographs from the Osteoarthritis Initiative dataset. Two convolutional neural network models (ResNet18 and ConvNeXt-Small) were trained to classify osteophytes, joint space narrowing, subchondral sclerosis, and Kellgren-Lawrence grades (KLG). These models were compared against two FM: BiomedCLIP, a multimodal vision-language model pre-trained on diverse medical images and text, and RAD-DINO vision transformer model pre-trained exclusively on chest radiographs. We extracted image embeddings from both FMs and used XGBoost classifiers to perform downstream classification. Performance was assessed using a comprehensive classification metrics appropriate for binary and multi-class classification tasks. DL models outperformed FM-based approaches across all tasks. ConvNeXt achieved the highest performance in predicting KLG, with a weighted Cohen's kappa of 0.880 and higher AUC in binary tasks. BiomedCLIP and RAD-DINO performed similarly, and BiomedCLIP's prior exposure to knee radiographs during pretraining led to only slight improvements. Zero-shot classification using BiomedCLIP correctly identified 91.14% of knee radiographs, with most failures associated with low image quality. Grad-CAM visualizations revealed DL models, particularly ConvNeXt, reliably focused on clinically relevant regions. While FMs offer promising utility in auxiliary imaging tasks, supervised DL remains superior for fine-grained radiographic feature classification in domains with limited pretraining representation, such as musculoskeletal imaging.

Authors

  • Mohammadreza Chavoshi
    Department of Radiology, Tehran University of Medical Sciences, Shariati Hospital, Tehran, Iran.
  • Hari Trivedi
    Department of Radiology, Medical College of Georgia at Augusta University, 1120 15th St, Augusta, GA 30912 (Y.T.); and Department of Radiology, Emory University, Atlanta, Ga (B.V., E.K., A.P., J.G., N.S., H.T.).
  • Janice Newsome
    Division of Interventional Radiology and Image-Guided Medicine, Department of Radiology and Imaging Science, Emory University School of Medicine, Atlanta, Georgia.
  • Aawez Mansuri
    Department of Radiology, Emory University, Atlanta, GA, USA.
  • Frank Li
    Roy J. Carver Department of Biomedical Engineering, University of Iowa, Iowa City, IA 52242, USA.
  • Theo Dapamede
    School of Medicine, Emory University.
  • Bardia Khosravi
    Department of Radiology, Radiology Informatics Lab, Mayo Clinic, Rochester, MN 55905, United States.
  • Judy Gichoya
    Department of Radiology, Medical College of Georgia at Augusta University, 1120 15th St, Augusta, GA 30912 (Y.T.); and Department of Radiology, Emory University, Atlanta, Ga (B.V., E.K., A.P., J.G., N.S., H.T.).

Keywords

No keywords available for this article.