Benchmarking vision-language models for diagnostics in emergency and critical care settings.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 40640347.
- Also identified by DOI 10.1038/s41746-025-01837-2 and PMC identifier 12246445.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
The applicability of vision-language models (VLMs) for acute care in emergency and intensive care units remains underexplored. Using a multimodal dataset of diagnostic questions involving medical images and clinical context, we benchmarked several small open-source VLMs against GPT-4o. While open models demonstrated limited diagnostic accuracy (up to 40.4%), GPT-4o significantly outperformed them (68.1%). Findings highlight the need for specialized training and optimization to improve open-source VLMs for acute care applications.