Evaluating foundation models on official German medical licensing examinations: Implications for high-stakes assessment and AI-assisted medical education.

Cirkel, Lasse; Knitza, Johannes; Schillings, Volker; Oksche, Alexander; Becker, Jan Carl; Kuhn, Sebastian · NPJ Digit Med · 2026

Where this comes from

Abstract

We evaluated proprietary and open-weight foundation models on 24 German medical licensing examinations (2019-2024), including 7485 items and response data from 119,878 sittings. For fair comparison, the eight vision-capable models were evaluated on the full benchmark and all thirteen on a shared text-only subset. On the full benchmark, Gemini 3.1 Pro achieved the highest overall accuracy, reaching 99.31% on the first (M1) and 98.37% on the second (M2) examination. On the shared text-only subset, proprietary frontier models performed at near-ceiling levels, several open-weight models (including GLM-5 and DeepSeek V3.2-Thinking) were highly competitive, and even compact ones exceeded mean student performance. Image-present items were more difficult for both students and models, but the associated decline was disproportionately larger for models than for students. Human- and model-defined difficulty subsets showed limited overlap, and model-hard subsets revealed residual differences among top systems. These findings underscore the rapid progress of these models, particularly open-weight systems, and the value of official German medical licensing examinations as a restricted-access benchmark with reduced public exposure. They carry implications for high-stakes assessment and AI-assisted medical education, notably multimodal assessment, human-aligned educational tools, and privacy-preserving local deployment.