CVJun 12

Learning Sparse Latent Predictive Foundation Model for Multimodal Neuroimaging

arXiv:2606.1495714.1
Predicted impact top 30% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and clinicians working with multimodal brain MRI, Neuro-JEPA provides a scalable framework for learning unified representations that outperform existing methods across diverse clinical and research tasks.

Neuro-JEPA, a sparse multimodal neuroimaging foundation model combining latent predictive objective with Mixture-of-Experts, was pretrained on 1.55M scans and evaluated on 47 tasks across three health systems and 12 public datasets, achieving stronger and more consistent performance than existing neuroimaging foundation models and a simple CNN baseline.

Brain MRIs are routinely acquired as multiple complementary sequences with unique contrast weighting, including T1-weighed imaging (T1w) anatomic and fluid-sensitive T2-weighted (T2w) contrasts. However, methods for learning unified representations across the multitude of MRI contrast mechanisms at health-system scale are lacking. In this study, we introduce Neuro-JEPA, a sparse multimodal neuroimaging foundation model that combines a latent predictive objective with a Mixture-of-Experts architecture to encode brain MRI across core T1w, T2w, and fluid-suppressed FLAIR imaging (FLAIR). We further provide a systematic methodological study of architectural, masking, objective, and sparsity design choices beneficial for robust neuroimaging multimodal representation learning. Neuro-JEPA was pretrained on 1,551,862 scans from 428,647 studies after modality-specific preprocessing with data curation across three core structural brain MRI sequences. We evaluated the learned representations across clinical and research settings, including 25 tasks from three health systems: NYU Langone, NYU Long Island, and Massachusetts General Hospital, and 22 tasks from 12 public datasets, covering unimodal, multimodal and cross-domain evaluation configurations. Across these benchmarks, existing neuroimaging foundation models showed inconsistent gains over a simple convolutional neural network (CNN) baseline, whereas Neuro-JEPA achieved stronger and more consistent performance across all evaluated settings. These results establish a scalable methodological framework for multimodal neuroimaging representation learning and highlight the need for foundation model evaluation protocols that include simple baselines, clinically heterogeneous cohorts and controlled multimodal comparisons.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes