Evaluating the Performance of Large Language Models on Palliative Care Test Questions: A Mixed Methods Study.

Chua, Isaac S; Lo, Yen-Ting; Liu, David; Succi, Marc D; Zhang, Mark; Yeh, Jonathan; Skarf, Lara M; Doyle, Kathleen et al. · J Palliat Med · 2026

other

Where this comes from

Abstract

Little is known about large language model (LLM) performance on palliative care (PC)-related knowledge-based tasks. We evaluated two LLMs in answering PC-related test questions and explaining their answer choice rationale. LLMs were prompted to answer 25 randomly selected questions from the <i>Fast Facts Quiz</i> and provide their answer choice rationale. Three PC educators ranked and rated LLM-generated answer choice explanations versus the test's answer key explanations. Linear fixed-effect models evaluated reviewer ranking, and ordinal logistic regression evaluated reviewer ratings of quality, suitability, accuracy, relevance, and comprehensiveness. Both LLMs answered 96% of selected questions correctly. Reviewers rated LLM-generated explanations more highly than <i>Fast Facts Quiz</i> explanations. Five themes emerged from reviewer comments: perceived inaccuracies, clarity of writing, educational value, linguistic style, and miscellaneous. LLMs demonstrated high answer choice accuracy and generated preferable answer explanations when compared to the <i>Fast Facts Quiz</i> answer key.

Medical subject headings