Multilingual Hematology Visual Question Answering Dataset
Abstract
Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understand- ing visual and textual information for tasks such as Visual Question Answering (VQA). However, existing hematology vision-language resources remain predominantly English-centric, limiting their applicability in multilingual healthcare environments. This challenge is particularly relevant in South Asia, especially in Pakistan, where Urdu is widely spo- ken, while healthcare information and digital healthcare systems remain largely English-based. To investigate this gap, we surveyed healthcare professionals and identified substantial language mismatches between clinical documentation and patient communication, underscoring the need for multilingual healthcare technologies. To address this limitation, we introduce WBCMor-VQA, a clinically validated bilingual (English–Urdu), morphology-aware VQA benchmark for leukemia and normal white blood cell (WBC) analysis. The benchmark is constructed using morphology-aware annota- tions from the LeukemiaAttri and WBCAtt datasets and is supported by a domain-specific Urdu hematology dictionary to ensure linguistic consistency and clinical correctness. The final benchmark comprises 110K bilingual question–answer pairs corresponding to 20K leukemic and normal single-cell images. Furthermore, we establish strong baseline results by evaluating multiple open-source VLMs on the proposed benchmark. The proposed resource aims to facilitate the development of accessible and clinically relevant AI systems for multilingual healthcare environments. The dataset is publicly available at: https://doi.org/10.6084/m9.figshare.32727159
// Source
Authors: Hajra Malik, Hafiza Tooba Aftab, Abdul Rehman, Mohsen Ali, Waqas Sultani
Institutions: King Edward Medical University, Information Technology University