A generalist biomedical vision-language model via multi-CLIP knowledge distillation.

Wang, Shansong; Jin, Zhecheng; Hu, Mingzhe; Safari, Mojtaba; Gao, Yuan; Zhao, Feng; Chang, Chih-Wei; Qiu, Richard Lj et al. · Nat Commun · 2026

basic_science · Level V

Where this comes from

Abstract

Contrastive Language-Image Pretraining (CLIP) models, which are pretrained on natural images with billions of image-text pairs, exhibit strong zero-shot and cross-modal capabilities. However, their application in biomedicine remains challenging due to limited large-scale image-text data and heterogeneous imaging modalities. Here we show that a generalist biomedical foundation model can be effectively built via multimodal medical knowledge distillation. We introduce MMKD-CLIP, which integrates complementary knowledge from nine biomedical CLIP models. Our two-stage pipeline combines CLIP-style pretraining on 2.9 million biomedical image-text pairs across 26 modalities with large-scale feature-level distillation. We evaluate MMKD-CLIP on 58 datasets spanning nine modalities and six tasks, including classification, retrieval, visual question answering, survival prediction, and cancer diagnosis. MMKD-CLIP performs favorably relative to the teacher models, with results supporting its robustness and cross-domain generalization.