TMF-RSE: Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty for Lung Severity Scoring
Provides accurate severity quantification with uncertainty estimates for clinical decision-making in lung disease assessment.
TMF-RSE introduces a tri-modal fusion framework combining appearance, structural, and semantic features with evidential uncertainty for lung severity scoring, achieving MAE of 4.02 and Pearson correlation of 0.9629 on Per-COVID-19 CT and 0.339 MAE / 0.973 PC on RALO geographic extent, outperforming transformer baselines.
Accurate quantification of lung disease severity from chest imaging is critical for clinical decision-making and resource allocation. We propose a tri-modal deep learning framework, TMF-RSE (Tri-Modal Fusion with Regional Semantics and Evidential Uncertainty), that combines appearance features from two-dimensional chest inputs, structural features from lung segmentation masks, and semantic features from vision-language models (VLMs) for severity quantification. Our approach employs complementary fusion mechanisms that integrate semantic guidance, structural priors, and hierarchical interactions across modalities. The model employs evidential regression to provide both severity predictions and uncertainty estimates. Experiments on the Per-COVID-19 CT and RALO datasets show that TMF-RSE outperforms recent transformer-based baselines, achieving MAE of 4.02 and Pearson correlation of 0.9629 on Per-COVID-19 validation, and 0.339 MAE / 0.973 PC on RALO geographic extent.