Evaluating the role of large language models in traditional Chinese medicine diagnosis and treatment recommendations.
cross_sectional · Level IV
Where this comes from
- Record sourced from PubMed, PMID 40691277.
- Also identified by DOI 10.1038/s41746-025-01845-2 and PMC identifier 12279949.
- Licence recorded as CC BY-NC-ND.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Digital health technologies hold significant potential for reducing global healthcare disparities. Large language models (LLMs) offer new opportunities to enhance access to culturally specific healthcare, including traditional Chinese medicine (TCM). This study evaluated the diagnostic and treatment performance of seven publicly available LLMs using a real-world acupuncture case, comparing their outputs with three professional acupuncturists across five domains: Western diagnosis, TCM diagnosis, acupoint selection, needling technique, and herbal medicine. Twenty-eight expert evaluators from China, South Korea, and the United States assessed the responses using a multilingual survey. LLMs performed comparably to acupuncturists in Western diagnosis and showed variable performance in TCM-specific tasks. GPT-4o, Qwen 2.5 Max, and Doubao 1.5 Pro demonstrated the highest alignment with expert evaluations, particularly in TCM diagnosis and acupoint selection. These findings highlight the potential of general-purpose LLMs to support culturally grounded medical decision-making and reduce access barriers in TCM care systems.