Utilizing Data Augmentation to Assist Melanoma Detection Models Across Diverse Racial Groups

Authors

  • Nam Tran East Hamilton High School

Abstract

Current machine learning (ML) models developed to tackle the issue of melanoma detection struggle to detect melanoma in diverse skin groups, namely individuals from African and Hispanic backgrounds. This problem primarily stems from a lack of diverse data available to train ML models. As a solution, this paper introduces the use of data augmentation to correct both a lack of data diversity and class imbalances. For this study, data augmentation was applied to better enhance the performance of two ML models for diverse melanoma detection: ResNet50 and Random Forest. Using the state-of-the-art HAM10000 and Diverse Dermatology Images (DDI) datasets, the researcher examined how well each model did in distinguishing melanoma from both the lighter skin in the HAM10000 dataset and the darker skin in the DDI dataset before and after data augmentation was applied. When assessed, both models across both datasets saw improvements across all metrics, especially recall. Overall, however, ResNet50 (accuracy = 0.55) and Random Forest (accuracy = 0.44) obtained significantly worse results when shown the darker-skinned images. Thus, while the results reaffirm the practice of data augmentation to correct poor class balances and to increase the diversity of training data, its use was not able to reduce the disparity gap between Caucasian and darker racial groups in terms of melanoma detection.

Downloads

Published

2026-08-24

Data Availability Statement

The data is available in the Abstract. Additional data may be accessed upon request.

Issue

Section

Research Articles