CVJul 9

CT-CLIP Representations for Multimodal Lung Cancer Survival Prediction

arXiv:2607.085034.5h-index: 8
Predicted impact top 79% in CV · last 90 daysOriginality Synthesis-oriented
AI Analysis

For clinicians and researchers in lung cancer prognosis, this work demonstrates that a frozen foundation model can provide accurate survival predictions in data-constrained settings, reducing the need for large curated imaging datasets.

The study evaluates CT-CLIP, a domain-specific foundation model, for multimodal lung cancer survival prediction using CT images and clinical data from 242 patients. A frozen CT-CLIP with a lightweight survival head outperforms clinical baselines and achieves comparable or better performance than other multimodal methods, effectively stratifying patients into risk groups.

Accurate prognosis prediction is important for treatment planning in lung cancer, but deep learning-driven survival modelling is often limited by the scarcity of curated imaging cohorts with reliable outcome data. This study evaluates whether representations from a domain-specific foundation model can be used for multimodal survival prediction in data-constrained clinical settings. We assess the foundation model CT-CLIP as a feature extractor for pretreatment computed tomography images and clinical variables from 242 diagnosed lung cancer patients. The evaluation includes adaptation strategies based on frozen encoders, full fine-tuning, and low-rank adaptation, together with modality ablations and comparisons with clinical and multimodal baselines. The results show that a frozen CT-CLIP model combined with a trainable lightweight survival head outperforms the clinical baseline and achieves comparable or improved performance relative to other multimodal approaches, and separates patients into clinically meaningful high- and low-risk groups.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes