D<sup>3</sup>-KD: Breast cancer segmentation via data diffusion-enhanced knowledge distillation under source-free domain conditions.

Cai, Jiati; Wang, Kun; Zhong, Meihui; Zhang, Haohan; Tong, Yuxin; Wang, Junren; Zhou, Fan; Jing, Jing et al. · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

In breast ultrasound segmentation, there is a strong demand for lightweight yet high-quality models for practical deployment. Knowledge distillation (KD) is a natural solution for transferring knowledge from high-capacity models to compact ones. However, the teacher's training domain is often inaccessible in clinical practice, leading to a source-domain-unknown KD setting. Under this condition, domain discrepancies between teacher and student distributions can significantly reduce distillation effectiveness. From a theoretical perspective, we analyze KD under domain discrepancies and show that, while KD can still be beneficial, its gains diminish when the support overlap between the teacher and target distributions is small. Our analysis further suggests that enlarging the support of the target distribution can increase this overlap and thereby strengthen knowledge transfer. Motivated by this insight, we propose a Data Diffusion-Enhanced Knowledge Distillation paradigm (D<sup>3</sup>-KD) that augments target-domain data to facilitate knowledge transfer under unknown source domains. Specifically, leveraging the generative capability of diffusion models, we design a Drift-Corrected Efficient Diffusion (DCED) module to effectively expand the target data distribution while correcting undesired drift and improving sampling efficiency via shifted initialization. Extensive experiments on breast ultrasound segmentation demonstrate that D<sup>3</sup>-KD consistently improves distillation performance in source-domain-unknown scenarios.