Performance of Three Conversational Artificial Intelligence Agents in Defining End-of-Life Care Terms.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 40138176.
- Also identified by DOI 10.1089/jpm.2024.0526.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
<b><i>Background:</i></b> Conversational artificial intelligence agents, or chatbots, are a transformational technology understudied in end-of-life care. <b><i>Methods:</i></b> OpenAI's ChatGPT, Google's Bard, and Microsoft's Bing were asked to define "terminally ill," "end of life," "transitions of care," "actively dying," and provide three references. Outputs were scored by six physicians on a scale of 0-10 for accuracy, comprehensiveness, and credibility. Flesch-Kincaid Grade Level and Flesch Reading Ease (FRE) were used to calculate readability. <b><i>Results:</i></b> Mean (standard deviation) scores for accuracy were 9 (1.9) for ChatGPT, 7.5 (2.4) for Bard, and 8.3 (2.4) for Bing. Comprehensiveness scores averaged 8.5 (1.7) for ChatGPT, 7.3 (2.1) for Bard, and 6.5 (2.3) for Bing. Credibility was low with a mean score of 3 (1.8). The mean FRE score was 41.7, and the mean grade level was 14.1, indicating low readability. <b><i>Conclusion:</i></b> Chatbot outputs had important deficiencies that necessitated clinician oversight to prevent misinformation.
Medical subject headings
- Artificial Intelligence
- Terminal Care
- Communication
- Terminology as Topic