Multi-Agent collaboration as a complementary architecture for AI-generated medical examination items.

Jiang, Zhehan · NPJ Digit Med · 2026

other · Level V

Where this comes from

Abstract

Qian et al. showed single LLMs can generate acceptable knowledge-based questions but struggle with higher-order reasoning. We argue this is architectural: decomposing item development into specialized agents for drafting, critique, and iterative adversarial refinement improves quality. In blinded evaluation for China's National Medical Licensing Examination, multi-agent outputs received 57.7% of expert preferences, compared with 42.3% for the single-model baseline, indicating a viable path to surpass current limits in AI-assisted item generation.