Evaluating the Performance of Large Language Models on Palliative Care Test Questions: A Mixed Methods Study.
other
Where this comes from
- Record sourced from PubMed, PMID 42231120.
- Also identified by DOI 10.1177/10966218261456817.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Little is known about large language model (LLM) performance on palliative care (PC)-related knowledge-based tasks. We evaluated two LLMs in answering PC-related test questions and explaining their answer choice rationale. LLMs were prompted to answer 25 randomly selected questions from the <i>Fast Facts Quiz</i> and provide their answer choice rationale. Three PC educators ranked and rated LLM-generated answer choice explanations versus the test's answer key explanations. Linear fixed-effect models evaluated reviewer ranking, and ordinal logistic regression evaluated reviewer ratings of quality, suitability, accuracy, relevance, and comprehensiveness. Both LLMs answered 96% of selected questions correctly. Reviewers rated LLM-generated explanations more highly than <i>Fast Facts Quiz</i> explanations. Five themes emerged from reviewer comments: perceived inaccuracies, clarity of writing, educational value, linguistic style, and miscellaneous. LLMs demonstrated high answer choice accuracy and generated preferable answer explanations when compared to the <i>Fast Facts Quiz</i> answer key.
Medical subject headings
- Large Language Models
- Palliative Care
- Educational Measurement