Benchmarking the symptom-checking capabilities of ChatGPT for a broad range of diseases.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 38109889.
- Also identified by DOI 10.1093/jamia/ocad245 and PMC identifier 11339504.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
This study evaluates ChatGPT's symptom-checking accuracy across a broad range of diseases using the Mayo Clinic Symptom Checker patient service as a benchmark. We prompted ChatGPT with symptoms of 194 distinct diseases. By comparing its predictions with expectations, we calculated a relative comparative score (RCS) to gauge accuracy. ChatGPT's GPT-4 model achieved an average RCS of 78.8%, outperforming the GPT-3.5-turbo by 10.5%. Some specialties scored above 90%. The test set, although extensive, was not exhaustive. Future studies should include a more comprehensive disease spectrum. ChatGPT exhibits high accuracy in symptom checking for a broad range of diseases, showcasing its potential as a medical training tool in learning health systems to enhance care quality and address health disparities.
Medical subject headings
- Benchmarking