CVJul 6

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

arXiv:2607.0469412.7
Predicted impact top 26% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For researchers and practitioners applying VLMs to medical AI, this work identifies and benchmarks a previously overlooked but essential preprocessing step.

This paper introduces the Medical Data Standardization Benchmark (MDS-Bench) to evaluate VLMs on raw medical data standardization, finding that even the best model (Gemini 3 Flash) achieves only 48.6% end-to-end success rate, highlighting a critical bottleneck for real-world medical AI.

As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources. In this paper, we study this missing step, i.e., raw medical data standardization. Specifically, models are given raw dataset folders and evaluated on their ability to identify source formats, convert raw medical images into VLM-compatible visual inputs, extract relevant textual information, and organize the results into structured image-text pairs. To construct this Medical Data Standardization Benchmark (MDS-Bench), we manually annotate 1,939 raw medical data standardization tasks covering diverse clinical practice, radiology modalities, annotation formats, and directory layouts. Extensive experiments show that even the best performing VLMs, i.e., Gemini 3 Flash, achieve only 48.6% end-to-end success rate. Our research highlights raw medical data standardization as a critical bottleneck for medical AI diagnosis in real practice.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes