Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41150678.
- Also identified by DOI 10.1371/journal.pone.0335775 and PMC identifier 12561936.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Multimodal large-language models (LLMs) have recently gained the ability to interpret images. However, their accuracy on anatomy tasks remains unclear. A cross-sectional, atlas-based benchmark study was conducted in which six publicly accessible chat endpoints, including paired "deep-reasoning" and "low-latency" modes from OpenAI, Microsoft Copilot, and Google Gemini, identified 260 numbered landmarks on 26 high-resolution plates from a classical anatomic atlas. Each image was processed twice per model. Two blinded anatomy lecturers scored responses, including accuracy, run-to-run consistency, and per-label latency, which were compared with χ² and Kruskal-Wallis tests. Overall accuracy differed significantly among models (χ² = 73.2, P < 0.001). OpenAI o3 achieved the highest correctness (53.1%), outperforming its sibling GPT-4o and both Copilot variants, but required the longest inference time. Musculoskeletal structures were recognised more accurately than neurovascular targets, reflecting the greater visual complexity of fine vessels and nerves. Consistency ranged from 43.5% (Gemini Flash) to 65.0% (GPT-4o); deeper modes improved stability for Copilot and Gemini but not accuracy. Median per-label latency spanned three orders of magnitude, from 0.5 s for Gemini Flash to 33 s for o3. Currently, publicly available multimodal LLMs can only moderately identify oral and maxillofacial landmarks, and no endpoint is sufficiently reliable to serve as a stand-alone answer key. Higher accuracy was achievable with a trade-off in latency, highlighting the need for domain-specific tuning and human oversight. This atlas benchmark study introduced here provides a reproducible yardstick for future model refinement and educational integration.
Medical subject headings
- Benchmarking
- Language
- Models, Anatomic
- Mouth
- Face