Abstract

Cutaneous melanoma is one of the most aggressive skin cancers, where early detection is essential to improve survival rates and reduce invasive treatments. Although convolutional neural networks (CNNs) have shown outstanding performance in dermoscopic image classification, their limited interpretability raises concerns about their reliability in clinical practice. This study aims to analyze the Clever Hans effect in CNN-based binary melanoma classification, focusing on the influence of spurious visual artifacts and the impact of preprocessing in mitigating such biases. To the best of our knowledge, this is the first study to systematically investigate this phenomenon in the context of binary classification of dermoscopic images, highlighting the potential risk of CNNs relying on non-lesion-related features. Seven widely used CNN architectures (ResNet-50, ResNet-101, VGG16, VGG19, DenseNet-121, MobileNetV2, and EfficientNet-B0) were trained on four public dermoscopic datasets (ISIC 2016, PH2, ISIC 2020, and HAM10000), both with and without a preprocessing pipeline including artifact removal (e.g., hair, ink, scale bars). Model interpretability was assessed using seven Grad-CAM variants: Grad-CAM, Grad-CAM++, Score-CAM, Ablation-CAM, Layer-CAM, Eigen-CAM, and XGrad-CAM, which were quantitatively evaluated using Dice coefficient and Intersection over Union (IoU) on ISIC 2016 and PH2, which contain expert segmentation masks. Classification performance remained consistently high. In ISIC 2016, ResNet-101 yielded the best results without preprocessing (87.2% accuracy, 0.871 AUC-ROC), while MobileNetV2 performed best after preprocessing (84.5%, 0.850). In PH2 and HAM10000, DenseNet-121 remained the best model in both settings, though metrics slightly decreased post-preprocessing (e.g., accuracy dropped from 92.3% to 88.5% in HAM10000). In ISIC 2020, DenseNet-121 achieved top performance without preprocessing, while ResNet-50 led after preprocessing. Interpretability improved notably: in ISIC 2016, Dice increased from 0.356 to 0.604 and IoU from 0.242 to 0.424; in PH2, Dice rose from 0.367 to 0.716 and IoU from 0.254 to 0.564. These results underscore that, while preprocessing may lead to slight reductions in classification metrics, it substantially improves the model¿s ability to focus on lesion areas and reduces its dependence on irrelevant artifacts. Therefore, combining robust preprocessing with post hoc interpretability techniques is essential to develop reliable and clinically transparent AI systems in dermatology.
Loading...

Quotes

plumx
0 citations in WOS
0 citations in

Journal Title

Journal ISSN

Volume Title

Publisher

Universidad Rey Juan Carlos

URL external

External URL

DOI

Description

Trabajo Fin de Grado leído en la Universidad Rey Juan Carlos en el curso académico 2024/2025. Directores/as: Cristina Soguero Ruíz, Vanesa Gómez Martínez

Citation

Endorsement

Review

Supplemented By

Referenced By

Statistics

Views
3
Downloads
0

Bibliographic managers