Assessing the Usability of ChatGPT Responses Compared to Other Online Information in Hand Surgery.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 40219840.
- Also identified by DOI 10.1177/15589447251329584 and PMC identifier 11993548.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
ChatGPT is a natural language processing tool with potential to increase accessibility of health information. This study aimed to: (1) assess usability of online medical information for hand surgery topics; and (2) evaluate the influence of medical consensus. Three phrases were posed 20 times each to Google, ChatGPT-3.5, and ChatGPT-4.0: "What is the cause of carpal tunnel syndrome?" (high consensus), "What is the cause of tennis elbow?" (moderate consensus), and "Platelet-rich plasma for thumb arthritis?" (low consensus). Readability was assessed by grade level while reliability and accuracy were scored based on predetermined rubrics. Scores were compared via Mann-Whitney <i>U</i> tests with alpha set to .05. Google responses had superior readability for moderate-high consensus topics (<i>P</i> < .0001) with an average eighth-grade reading level compared to college sophomore level for ChatGPT. Low consensus topics had poor readability throughout. ChatGPT-4 responses had similar reliability but significantly inferior readability to ChatGPT-3.5 for low medical consensus topics (<i>P</i> < .01). There was no significant difference in accuracy between sources. ChatGPT-4 and Google had differing coverage of cause of disease (<i>P</i> < .05) and procedure details/efficacy/alternatives (<i>P</i> < .05) with similar coverage of anatomy and pathophysiology. Compared to Google, ChatGPT does not provide readable responses when providing reliable medical information. While patients can modulate ChatGPT readability with prompt engineering, this requires insight into their health literacy and is an additional barrier to accessing medical information. Medical consensus influences usability of online medical information for both Google and ChatGPT. Providers should remain aware of ChatGPT limitations in distributing medical information.
Anatomy
- hand