Controllable Echocardiogram Video Generation Driven by Fine-Grained Cardiac Motion.

Wang, Jinduo; Lu, Zhi; Wang, Binquan; Zhang, Dongheng; Hu, Yang; Zhu, Da; Pan, Xiangbin; Chen, Yan · IEEE Trans Med Imaging · 2026

basic_science · Level V

Where this comes from

Abstract

Echocardiogram video synthesis has emerged as a promising solution to alleviate data scarcity for training intelligent diagnostic models and to enhance clinical education. However, existing methods are typically conditioned on a single scalar metric, such as left ventricular ejection fraction (LVEF), which fails to capture the complex temporal dynamics of cardiac motion and thus limits clinical applicability. To address this limitation, we propose a novel image-to-video synthesis framework guided by left ventricular volume-time curves (VTCs), which provide a more comprehensive representation of cardiac function beyond conventional scalar indicators. Specifically, we develop a VTC-conditioned diffusion model for controllable echocardiogram video generation. Building upon this formulation, we introduce a two-stage architecture, where a latent optical flow module captures motion dynamics, and a conditioned image-to-flow sequence model subsequently generates temporally coherent motion patterns. This design enables efficient echocardiographic video generation with physiologically consistent cardiac dynamics, while the adversarial loss further improves visual fidelity by reducing blurring artifacts. In addition, to address the scarcity of labeled data, we propose a semi-supervised framework for extracting VTCs directly from echocardiogram videos. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance, enabling fine-grained control of cardiac dynamics and advancing the practical applicability of generative models in echocardiography.