Multi-model Artificial Intelligence Evaluation in Sudden Sensorineural Hearing Loss.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41603577.
- Also identified by DOI 10.1002/ohn.70143.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
To compare the diagnostic accuracy, linguistic clarity, and user satisfaction of three large language models (ChatGPT-4.0, Claude 3.7 Sonet, and OpenAI Mini 3) in managing sudden sensorineural hearing loss. Prospective, multi-domain comparative analysis using blinded expert evaluation. Online artificial intelligence (AI) platforms accessed under standardized conditions. Twenty-seven sudden sensorineural hearing loss-related questions-covering general knowledge, audiometric interpretation, and clinical case scenarios-were submitted to the three AI models. Responses were evaluated by 10 board-certified otolaryngologists using three validated tools: Quality Assessment of Medical Artificial Intelligence (QAMAI), Artificial Intelligence Performance Instrument (AIPI), and Artificial Intelligence Satisfaction and Performance Evaluation Questionnaire (AISPE-Q). Linguistic complexity was assessed using metrics such as word count, sentence length, lexical diversity, and clinical verb use. ChatGPT-4.0 demonstrated the highest scores in clinical accuracy (QAMAI: 4.57), completeness (4.53), and evaluator satisfaction (AISPE-Q: 94%). Claude 3.7 outperformed in clarity and sentence complexity, while OpenAI Mini 3 exhibited the highest lexical diversity and directive tone but scored lower overall. Inter-rater reliability was strong (intraclass correlation coefficient [ICC] > 0.85). Correlation analysis revealed a significant relationship between objective quality and subjective satisfaction (r > 0.76). ChatGPT-4.0 delivered the most clinically aligned and satisfactory responses, whereas Claude 3.7 provided linguistically refined outputs. Our findings support the context-specific application of hybrid large language model approaches in otolaryngology, particularly for patient education, diagnosis, and AI-driven triage. 2-prospective comparative diagnostic accuracy study.
Medical subject headings
- Hearing Loss, Sensorineural
- Artificial Intelligence
- Hearing Loss, Sudden