Automated production of comparison tables for shared decision making: Comparing a human-generated table (Option Grid), a search engine process, and outputs from four large language models.
other · Level V
Where this comes from
- Record sourced from PubMed, PMID 40987186.
- Also identified by DOI 10.1016/j.pec.2025.109356.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To explore the ability of artificial intelligence to produce comparison tables to facilitate shared decision-making. An expert human-generated comparison table (Option Grid ™) was compared to four comparison tables produced by large language models and one produced using a Google search process that a patient might undertake. Each table was prepared for a patient with osteoarthritis of the knee, considering a knee replacement. The results were compared to the Option Grid.™ The information items in each comparison table were divided into eight categories: the intervention process; benefits; side effects & adverse effects; pre-operative care; post-operative care & physical recovery; repeat surgery; decision-making process; and alternative interventions. We assessed the accuracy of each information item in a binary manner (accurate, inaccurate). OpenBioLLM-70b and two proprietary ChatGPT models generated similar frequencies of information items across most categories, but omitted information on alternative interventions. The Google search process yielded the highest number of information items (n = 41), and OpenBioLLM-8b yielded the lowest (n = 20). Accuracy, compared to the human Option Grid, was 97 % for the ChatGPT models and the open-source OpenBioLLM-70b, and 95 % for OpenBioLLM-8b and the Google search process. The human-generated Option Grid had superior readability. Large language models produced comparison tables that are 3-5 % less accurate than a human generated Option Grid. Comparison tables produced by large language models may be less readable and require additional checking and editing. Subject to fact-checking and feedback, large language models may have a role to play in scaling up the production of evidence-based comparison tables that could assist patients and others.
Medical subject headings
- Decision Making, Shared
- Search Engine
- Artificial Intelligence
- Language
- Decision Making