Evaluating skin tone scales for dermatologic dataset labeling: a prospective-comparative study.
prospective_cohort · Level II
Where this comes from
- Record sourced from PubMed, PMID 41429926.
- Also identified by DOI 10.1038/s41746-025-02245-2 and PMC identifier 12749783.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Skin tone affects artificial intelligence (AI) performance in dermatology. While labeling datasets for skin tone could improve algorithm generalizability for detecting dermatologic malignancies, large-scale validation of skin tone assessments is lacking. This prospective observational study assessed reliability of subjective tools (Fitzpatrick Skin Type [FST], Monk Skin Tone [MST], Pantone SkinTone Guide) and an objective colorimeter for in-person and photography-based settings to evaluate utility for labeling dermoscopic datasets. Colorimetry (gold standard for color measurement) demonstrated high precision with in-person measurements. Of subjective scales, MST demonstrated slightly tighter clustering in the color space and high repeatability for in-person and photography-based assessments (latter varied by lighting). Dermoscopic image-extracted color values correlated poorly with colorimetry values. For subjective ratings, MST more effectively captured differences in AI melanoma classification scores than FST. Findings underscore that FST is not a proxy for skin tone; an important role remains for skin tone assessment to improve AI performance.