Contrastive language image pretraining for a cardiac magnetic resonance image embedding with zero-shot capabilities.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42168185.
- Also identified by DOI 10.1038/s41467-026-73022-2.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Vision-language models trained using self-supervised learning are crucial to reduce the dependency on large volumes of labeled data. However, conventional self-supervised approaches that rely on precise image-text pairing are not always feasible for cardiovascular magnetic resonance imaging (CMR) given its ability to visualize cardiac anatomy, physiology, and microstructure in a single exam. We present CMR-contrastive language image pretraining (CMR-CLIP), a vision language model which treats CMR images as videos to jointly learn embeddings between the images in the study and associated reports. The model is trained on a large dataset consisting of 11,028 studies performed at a single healthcare institution and evaluated on an internal test (N = 2,758) and external dataset (N = 428). CMR-CLIP achieves remarkable performance in real-world clinical tasks, achieving accuracies of 88.5% for non-ischemic cardiomyopathy, 88.0% for ischemic cardiomyopathy, 96.2% for cardiac amyloidosis, and 98.6% for hypertrophic cardiomyopathy, potentially leading to more consistent diagnosis of cardiovascular disease.
Medical subject headings
- Magnetic Resonance Imaging
- Heart
- Image Processing, Computer-Assisted