A multimodal instruction dataset and benchmark for ultrasound understanding.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 42362710.
- Also identified by DOI 10.1038/s41746-026-02930-w.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
While promising in general medical imaging like CT and MRI, large vision-language models (LVLMs) struggle in ultrasonography due to a profound visual-semantic gap and a scarcity of instruction-following data. Translating the inherently noisy and operator-dependent ultrasound images into precise clinical descriptions remains a complex reasoning challenge for general-purpose vision-language models. To bridge this gap, we develop SonoInstruct, a large-scale multi-source dataset curated from over 30 ultrasound sources, providing 110k+ images and 260k+ instruction-following instances. For comprehensive evaluation, we establish SonoBench, a multi-dimensional benchmark suite that assesses eight core capabilities, including performance in out-of-distribution (OOD) scenarios. To verify the effectiveness of our dataset, we fine-tuned Qwen3-VL-2B-Instruct on SonoInstruct, yielding Qwen3-VL-2B-Sono, which achieves a 30.3% relative improvement over the base model on the SonoBench benchmark. These results demonstrate that SonoInstruct effectively bridges the gap between noisy sonographic images and clinical reasoning, establishing a robust foundation for AI-driven ultrasound applications.