Deep multimodal fusion of temporal bone CT and clinical variables for hearing outcomes after cholesteatoma surgery.

Koyama, Hajime; Kashio, Akinori; Kondo, Kenji · NPJ Digit Med · 2026

other · Level V

Where this comes from

Abstract

This single-center study developed an automated system for creating regions of interest on temporal bone CT images and three postoperative hearing prediction models, integrating CT images with clinical variables for cholesteatoma surgery, including tympanoplasty with or without mastoidectomy. Automated creation of a sphere-shaped region of interest without manual annotation achieved a 1.22-mm median center error and a 6.6-mm maximum distance error. A tabular model using only clinical variables, a residual model with image fusion, and a gated residual model that conservatively incorporated image information were evaluated. The mean absolute errors for postoperative air-conduction pure-tone average were 12.75 dB, 11.10 dB, and 9.91 dB, respectively, with a significant difference between the tabular and gated residual models. The corresponding areas under the curves for predicting a postoperative air-bone gap of ≤ 20 dB were 0.65, 0.69, and 0.71. This study provides proof-of-concept evidence that gated residual fusion of preoperative CT and clinical variables may improve preoperative hearing outcome prediction, although external validation and further development of clinically aligned surgical-planning frameworks are required before clinical use.