CVCLLGJun 18

Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology

arXiv:2606.2047712.9
Predicted impact top 34% in CV · last 90 daysOriginality Incremental advance
AI Analysis

For medical AI researchers, this work provides a scalable method to train spatially grounded VLMs without manual annotations, enabling verifiable outputs in radiology.

The authors introduce RefRad2D, a large-scale bilingual dataset of 1.2M CT/MR image-text pairs for radiology, and train RadGrounder, a VLM that jointly performs report generation, VQA, and spatial grounding. On external VQA benchmarks, RadGrounder achieves competitive results with specialized medical VLMs, and adding grounding supervision does not degrade language quality.

We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a large-scale bilingual (German/English) dataset of 1.2M CT and MR image-text pairs derived from clinical practice, with task-specific VQA and spatial grounding subsets generated automatically via LLM-based curation and automated segmentation. Trained on this data, our model RadGrounder jointly performs report generation, visual question answering, and spatial grounding via bounding-box detection or segmentation. On external VQA benchmarks (Slake, VQA-RAD), RadGrounder achieves competitive results with specialized medical VLMs. Adding our clinical data to the training mixture improves open-ended VQA over fine-tuning on the downstream datasets alone, showing the transferability of our dataset. Crucially, adding grounding supervision does not degrade language quality, enabling spatially verifiable outputs at no cost to VQA performance.

Foundations

The foundational work for this paper's niche, ranked by how specifically the neighbourhood builds on it — not by global fame.

Your Notes