Clinician-perceived clinical accuracy and guideline alignment of ChatGPT responses to patient questions on HPV and CIN: a two-year comparative evaluation.
Where this comes from
- Record sourced from PubMed, PMID 42537621.
- Also identified by DOI 10.1016/j.ijmedinf.2026.106636.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) such as ChatGPT are increasingly used by patients to obtain medical information, particularly for human papillomavirus (HPV) infection and cervical intraepithelial neoplasia (CIN), where patients seek explanations on diagnosis, risk, screening, and follow-up. The clinical usefulness of chatbot-generated information depends not only on the accuracy but also on its perceived relevance to clinical practice. This study evaluated perceived clinical accuracy and perceived alignment with clinical consensus of ChatGPT responses to common patient questions on HPV and CIN over time. In this longitudinal evaluation study, ten frequently asked patient questions on HPV and CIN were extracted from a Dutch cancer information platform (kanker.nl). Questions were submitted to the publicly available ChatGPT web interface at two timepoints (Aug-2024 and Oct-2025) using an identical procedure. Responses were assessed by five gynecologists using 5-point Likert scales for perceived clinical accuracy and perceived alignment with clinical consensus. Ratings of 4 or 5 were considered favorable evaluations. Outcomes were summarized descriptively, and inter-rater reliability was assessed using Fleiss' kappa. In 2024, 88.0 % of perceived clinical accuracy ratings and 82.0 % of perceived consensus alignment ratings were scored as 4 or 5. In 2025, these proportions were 92.0 % and 76.0 %, respectively. Favorable ratings were common for both outcomes at both timepoints. Modest differences were observed across individual questions, although no formal hypothesis testing was performed. Variability in expert ratings was observed, with Fleiss' kappa values below zero at both timepoints, indicating agreement worse than expected by chance. ChatGPT-generated responses to HPV- and CIN-related patient questions frequently received favorable clinician ratings. However, poor inter-rater reliability and the lack of objective guideline benchmarking limit the strength of these findings. Further validation is required before broader integration into patient education.