Clinical Applications of Multimodal Artificial Intelligence in Otolaryngology: A State-of-the-Art Review.
review · Level V
Where this comes from
- Record sourced from PubMed, PMID 42117403.
- Also identified by DOI 10.1002/ohn.70285 and PMC identifier 13418058.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Artificial intelligence (AI) has advanced to simultaneously process visual, auditory, and textual inputs, providing users with "multimodal" AI. Given the clinical integration potential of these tools, otolaryngologists must stay informed. This study reviews current literature on applications of multimodal AI in otolaryngology. The MEDLINE, EMBASE, SCOPUS, Cochrane Library, Web of Science, and CINAHL databases. Databases were searched from the date of inception to March 4, 2025, following Preferred Reporting Items for Systematic Reviews and Meta-analyses extension for scoping reviews (PRISMA-ScR) guidelines. Studies on any application of multimodal AI in otolaryngology were included. Forty-four studies were included, with 55% (24/44) published in 2024 and 18% (8/44) in 2025. Image and text were the most commonly combined modalities (80%, 35/44), with emerging combinations including video with vector data (2%,1/44) and omics with text and/or image (14%, 6/44). Head and neck cancer was the most common subspecialty of focus (75%, 33/44), followed by general ear, nose, and throat (ENT) (11%, 5/44). All studies applied the models for clinical education (9%, 4/44) or decision support (91%, 40/44), assessing performance in areas such as board-style examination performance (accuracy: 37%-86%) or disease classification and prognostication (area under the receiver operating characteristic curve [AUC] 0.65-0.96). However, most studies were limited to small, single-institution samples and lacked prospective validation. Model error, data set bias, and language limitations underscore the need for further refinement. The application of multimodal large language models (LLMs) in otolaryngology is rapidly expanding. Clinicians must understand both the capabilities and limitations of these systems. Rigorous validation and ethical oversight will be essential to ensure the safe, equitable, and effective adoption in otolaryngologic care.
Medical subject headings
- Artificial Intelligence
- Otolaryngology