Multimodal Skin Disease Classification Using Vision Transformers, Medical Captioning, and Metadata Fusion: An Analysis on the ISIC 2024 Dataset.

Journal: Biomedical physics & engineering express
Published Date:

Abstract

Skin cancer and dermatological diseases are among the most prevalent global health conditions, where early and accurate diagnosis is critical for improving patient outcomes. Although deep learning models have achieved strong performance in dermoscopic image classification, many existing approaches primarily rely on visual features and make limited use of complementary clinical metadata and language-based context routinely considered by dermatologists. Recent vision-language models (VLMs), including medical-domain adaptations such as MedCLIP, have begun to show promise in dermatology; however, their integration with structured clinical metadata and the impact of different multimodal fusion strategies have not been systematically analyzed. In this work, we address the binary skin lesion classification problem by conducting a structured evaluation of MedCLIP-based multimodal embeddings combined with classical machine learning and neural classifiers. Image-text representations are extracted using MedCLIP and fused with patient metadata through early and attention-based fusion mechanisms, followed by multilayer perceptron (MLP) and ensemble classifiers. Experiments are performed on a curated subset of the ISIC 2024 dataset comprising 1,600 training and 400 test dermoscopic images with associated metadata. The proposed multimodal approach achieves an accuracy of 96% (95.7% exact) with AUROC = 0.987, outperforming unimodal baselines and demonstrating the complementary value of language and metadata for skin lesion diagnosis. This study provides a comprehensive analysis of MedCLIP-based multimodal learning in dermatology and highlights the importance of fusion design in vision-language-metadata systems for computer-aided diagnosis.

Authors

Keywords

No keywords available for this article.