Controllable synthesis of dermoscopic images using diffusion models for enhanced computer aided diagnosis and detection.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42425051.
- Also identified by DOI 10.1016/j.media.2026.104191.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Computer Aided Diagnosis/Detection (CAD) systems for skin lesion analysis are challenged by limited and imbalanced dermoscopic datasets, necessitating the use of advanced data augmentation techniques. In this paper, we introduce DiDGen, an innovative method employing text-to-image Diffusion models for high-quality Dermoscopic image Generation to enhance CAD performance. Specifically, we propose a dynamic prompting framework, DermPrompt, that leverages large language models to produce attribute-rich text prompts, improving image generation quality. We further refine generation control by incorporating a novel region-aware fine-tuning approach to build visual-textual alignments and a training-free pipeline for synthesizing lesion-mask pairs. Extensive experiments reveal that our proposed method outperforms existing generative methods in image fidelity and diversity, with downstream classifiers and segmentation models showing average improvements of 2.32% in F1 score and 3.16% in IoU score-all achieved with a single finetuning process. This approach offers an efficient solution for augmenting dermoscopic datasets and advancing skin lesion diagnosis.