Multilingual Hematology Visual Question Answering Dataset
For healthcare AI researchers and practitioners in multilingual regions (especially South Asia), this work provides a first-of-its-kind bilingual hematology VQA benchmark, but the contribution is incremental as it primarily extends existing datasets with Urdu translations.
The paper introduces WBCMor VQA, a bilingual (English/Urdu) visual question answering benchmark for hematology, containing 110K QA pairs for 20K cell images. It addresses the language gap in multilingual healthcare by providing a clinically validated resource, with baseline evaluations showing current VLMs underperform on Urdu queries.
Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for tasks such as Visual Question Answering. However, existing hematology vision-language resources remain predominantly English centric, limiting their applicability in multilingual healthcare environments. This challenge is releveant generally to South Asia and specifically to Pakistan, where Urdu is widely used despite healthcare information and digital medical systems being largely dependent on English. To investigate this gap, we conducted a survey among healthcare professionals, which revealed substantial language mismatches between clinical documentation and patient communication, emphasizing the need for multilingual healthcare technologies. To address this limitation, we introduce WBCMor VQA, a clinically validated bilingual English, Urdu morphology aware VQA benchmark for leukemia and normal white blood cell analysis. The benchmark is constructed using morphology-aware annotations from LeukemiaAttri and WBCAtt datasets and supported by a domain specific Urdu hematology dictionary to ensure linguistic consistency and clinical correctness. The final benchmark contains 110K bilingual question answer pairs serving as VQA annotations for 20K leukemic and normal single-cell images. Furthermore, we establish baseline performance by evaluating multiple open-source VLMs on the proposed benchmark. The proposed resource aims to facilitate the development of accessible and clinically relevant AI systems for multilingual healthcare environments.