Controllable synthesis of dermoscopic images using diffusion models for enhanced computer aided diagnosis and detection.

Shentu, Junjie; Watson, Matthew; Moubayed, Noura Al · Med Image Anal · 2026

basic_science · Level V

Where this comes from

Abstract

Computer Aided Diagnosis/Detection (CAD) systems for skin lesion analysis are challenged by limited and imbalanced dermoscopic datasets, necessitating the use of advanced data augmentation techniques. In this paper, we introduce DiDGen, an innovative method employing text-to-image Diffusion models for high-quality Dermoscopic image Generation to enhance CAD performance. Specifically, we propose a dynamic prompting framework, DermPrompt, that leverages large language models to produce attribute-rich text prompts, improving image generation quality. We further refine generation control by incorporating a novel region-aware fine-tuning approach to build visual-textual alignments and a training-free pipeline for synthesizing lesion-mask pairs. Extensive experiments reveal that our proposed method outperforms existing generative methods in image fidelity and diversity, with downstream classifiers and segmentation models showing average improvements of 2.32% in F1 score and 3.16% in IoU score-all achieved with a single finetuning process. This approach offers an efficient solution for augmenting dermoscopic datasets and advancing skin lesion diagnosis.