Cross-Modal Iteration Distillation for Robust IHD Screening: The IDNet Framework and A New Benchmark
For researchers and clinicians in IHD screening, this work provides a reproducible benchmark and a fusion method that improves multimodal performance, though it is an incremental step over existing multimodal approaches.
IDNet, a multimodal framework with a Cross-Modal Distillation Aggregator, achieves robust IHD screening by fusing retinal images and clinical variables, outperforming baselines on a new 50,410-image benchmark from UK Biobank.
Color Fundus Photography (CFP) offers a low-cost and non-invasive route for ischemic heart disease (IHD) screening, but current studies are limited by scarce public benchmarks and ineffective fusion of retinal images with sparse clinical variables. We propose IDNet, a multimodal framework with a Cross-Modal Distillation Aggregator (CDA) that uses learnable queries to sequentially integrate left-eye, right-eye, and clinical features, mitigating the imbalance between high-dimensional visual features and low-dimensional tabular inputs. We also construct a reproducible UK Biobank benchmark with open-source curation and quality-control pipelines, yielding 50,410 images from 25,205 subjects. On this benchmark, IDNet outperforms image-only, clinical-only, and several multimodal baselines, and CDA consistently improves multiple visual encoders as a plug-in fusion module.