IV CVMay 13, 2025

VIViT: Variable-Input Vision Transformer Framework for 3D MR Image Segmentation

Badhan Kumar Das, Ajay Singh, Gengyan Zhao, Han Liu, Thomas J. Re, Dorin Comaniciu, Eli Gibson, Andreas Maier

arXiv:2505.08693v21 citationsh-index: 9MLMI@MICCAI

Originality Incremental advance

AI Analysis

This addresses the problem of handling heterogeneous MR data for medical imaging researchers, though it is incremental as it adapts existing transformer methods to a specific domain bottleneck.

The paper tackled the challenge of variable input contrasts in 3D MR image segmentation by proposing VIViT, a transformer-based framework for self-supervised pretraining and finetuning, achieving mean Dice scores of 0.624 for brain infarct and 0.883 for brain tumor segmentation.

Self-supervised pretrain techniques have been widely used to improve the downstream tasks' performance. However, real-world magnetic resonance (MR) studies usually consist of different sets of contrasts due to different acquisition protocols, which poses challenges for the current deep learning methods on large-scale pretrain and different downstream tasks with different input requirements, since these methods typically require a fixed set of input modalities or, contrasts. To address this challenge, we propose variable-input ViT (VIViT), a transformer-based framework designed for self-supervised pretraining and segmentation finetuning for variable contrasts in each study. With this ability, our approach can maximize the data availability in pretrain, and can transfer the learned knowledge from pretrain to downstream tasks despite variations in input requirements. We validate our method on brain infarct and brain tumor segmentation, where our method outperforms current CNN and ViT-based models with a mean Dice score of 0.624 and 0.883 respectively. These results highlight the efficacy of our design for better adaptability and performance on tasks with real-world heterogeneous MR data.

View on arXiv PDF

Similar