How to benchmark medical AI agents.

Ruhrberg Estévez, Silas; Ferber, Dyke; van der Schaar, Mihaela; Kather, Jakob Nikolas · PLoS Med · 2026

Level V

Where this comes from

Abstract

Medical artificial intelligence research is shifting from single-task models toward multimodal large language model-based agents for complex clinical workflows, requiring benchmarks that assess clinical reasoning, process safety, and resource stewardship rather than final outputs alone.

Medical subject headings