WoundcareVQA: A multilingual visual question answering benchmark dataset for wound care.

Yim, Wen-Wai; Ben Abacha, Asma; Doerning, Robert; Chen, Chia-Yu; Xu, Jiaying; Subbarao, Anita; Yu, Zixuan; Xia, Fei et al. · J Biomed Inform · 2025

other · Level V

Where this comes from

Abstract

Introduce the task of wound care multimodal multilingual visual question answering, provide baseline performances, and identify areas of future study. A dataset of wound care multimodal multilingual visual question answering (VQA) was created using consumer health questions asked online. Practicing US medical doctors were tasked with providing metadata and expert responses labels. Several instruct-enabled, multilingual visual question answering models (GPT-4o, Gemini-1.5-Pro, and Qwen-VL) were tested to benchmark performances. Finally, automatic evaluations were tested against domain expert response ratings. A multilingual dataset of 477 wound care cases, 768 responses, 748 images, 3k structured data labels, 1362 translation instances, and 10k judgments was constructed (https://osf.io/xsj5u/). Metadata scores ranged from 0.32-0.78 accuracy depending on classification type; response generation performances 0.06 BLEU, 0.66 BERTScore, 0.45 ROUGE-L in English and 0.12 BLEU, 0.69 BERTScore, and 0.50 ROUGE-L in Chinese. We construct and explore the tasks of multimodal, multilingual VQA. We hope the work here can inspire further research in wound care metadata classification, VQA response generation, and open response automatic evaluation.

Medical subject headings