Evaluating the robustness and readiness of large frontier models in health AI applications.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 42362863.
- Also identified by DOI 10.1038/s41591-026-04501-8.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large frontier models such as GPT-5 and Gemini have demonstrated remarkable performance in a wide range of health application benchmarks. However, underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal reasoning. Here we systematically apply and integrate a series of adversarial stress tests to assess the robustness of flagship models and health benchmarks. Our study reveals prevalent brittleness in the presence of simple adversarial transformations: leading systems can guess the correct answer even with key inputs removed yet may get confused by the slightest prompt alterations while fabricating convincing but flawed reasoning traces. Using clinician-guided rubrics, we demonstrate that popular health benchmarks vary widely in what they truly measure. Our study reveals considerable gaps between benchmark performance and the robustness evidence needed to support claims about multimodal medical reasoning in health applications.