Temporally Consistent and Controllable Video Generation of 2D Cine CMR via Latent Space Motion Modeling
It addresses data scarcity in cardiac MRI for developing data-driven models by generating high-fidelity, on-demand medical video data.
The paper proposes a text-to-video generative method for synthesizing temporally coherent and anatomically consistent cardiac cine CMR sequences, achieving an FID of 31.68 and a CLIP score of 31.04.
Cine cardiac magnetic resonance is the gold standard for assessing cardiac function, but the scarcity of public datasets limits the development of advanced data-driven models. To address this limitation, we propose a generative method for synthesizing temporally coherent and anatomically consistent cardiac sequences. Our text-to-video framework decouples cardiac spatial structure from temporal motion. First, a fine-tuned diffusion model synthesizes an initial frame from a clinical text prompt, controlling anatomical features. Then, a latent flow model conditioned on a cardiac phase embedding generates the complete cardiac motion, ensuring spatial consistency and temporal control. Our model generates anatomically and pathologically diverse sequences with high temporal coherence and strong fidelity to input prompts, achieving a FID of 31.68 for image realism and a CLIP score of 31.04 for text-image alignment. These experimental results highlight its potential to produce high-fidelity, on-demand medical data, offering a scalable solution to data scarcity.