ChatGPT-4o Mini Fabricates and Miscites Evidence for American Academy of Orthopaedic Surgeons Hip Fracture Clinical Practice Guidelines.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 41625387.
- Also identified by DOI 10.2106/JBJS.OA.25.00225 and PMC identifier 12854652.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Generative artificial intelligence (AI) large language model (LLM) chatbots, such as ChatGPT, are increasingly used to answer medical questions. This study sought to assess the accuracy and quality of evidence cited in ChatGPT-4o mini responses to questions pertaining to hip fracture care. Prompt questions regarding hip fracture management that aligned with each of the 19 recommendations published in the American Academy of Orthopaedic Surgeons (AAOS) Clinical Practice Guideline (CPG) for Management of Hip Fractures in Older Adults were posed to the ChatGPT-4o mini LLM asynchronously by 4 independent medical student graders. Three prompt variations were applied for each recommendation, reflecting the perspectives of a physician, a patient, and a general information seeker. Graders then requested from the LLM a reference list with PubMed Identifier (PMID) numbers supporting each recommendation. Accuracy and clarity of responses were assessed using a standard rubric for overlap with CPG citations, fabrications, and inaccurate citations. ChatGPT-4o mini returned 228 responses to prompts seeking advice on AAOS CPG hip management recommendations. 76.3% of responses were "accurate" to the CPG recommendation. 88.2% of responses received a clarity rating of "excellent". ChatGPT-4o mini provided 228 responses citing 2,556 publications when prompted for supporting evidence, of which 1.1% overlapped with AAOS CPG references, and 7.9% were fabricated. Of the publications cited by the LLM which exist in the PubMed index, 91.7% were given with incorrect authors, 91.5% incorrect titles, 91.4% incorrect pages, 91.0% incorrect PMIDs, 90.9% incorrect journals, 90.3% incorrect journal volumes, and 20.0% incorrect publication years. Responses for an AAOS CPG strong recommendation strength were significantly more likely to be "accurate" (p < 0.001), and responses for an AAOS CPG limited strength recommendation were significantly more likely to be "unsupported" (p < 0.001). ChatGPT-4o mini provided clear, moderately accurate responses with rampantly erroneous and occasionally fabricated citations to queries about hip fracture care derived from the AAOS Clinical Practice Guideline on Management of Hip Fractures in Older Adults. Level V Therapeutic. See Instructions for Authors for a complete description of levels of evidence.
Anatomy
- hip