Explainable multimodal deep learning models for variable-length sequences in critically ill patients.

Martin, Jennifer; Afshar, Majid; Afshar, Askar Safipour; Caskey, John; Dligach, Dmitriy; Gao, Yanjun; Gao, Jifan; Chen, Guanhua et al. · J Biomed Inform · 2026

basic_science · Level V

Where this comes from

Abstract

Deep learning models have shown strong performance in predicting clinical events in critical care using structured electronic health record (EHR) data. While incorporating unstructured notes improves accuracy, multimodal fusion and explainability remain an open challenge, particularly for variable-length temporal data. This study develops an explainable temporal modeling framework for multimodal EHR data that accommodates variable-length intensive care unit (ICU) trajectories and supports diverse outcome prediction tasks. We introduced two multimodal recurrent neural networks (RNNs) with distinct fusion architectures (Pre-RNN and Post-RNN) that integrated structured EHR variables and unstructured clinical notes at every hourly timestep. Both architectures encoded temporal dynamics using Time2Vec and RNN layers with masking to handle variable-length sequences across patient stays. Models were benchmarked on four outcomes: 24-hour mortality, seven-day discharge, and four-hour ventilator or vasopressor onset in a publicly available EHR dataset. To enhance interpretability, integrated gradients was applied to estimate feature contributions from both modalities across timesteps, quantifying temporal and cross-modal importance. Multimodal fusion models outperformed unimodal baselines across all tasks, with Pre-RNN fusion achieving the highest area under the precision-recall curve (AUPRC) in three of four outcomes. Performance gains were modest for short-horizon events (ΔAUPRC < 0.01) but larger for intermediate and long-horizon (≥24 h) tasks. Integrated gradients revealed distinct attribution patterns, linking physiologic features (e.g., oxygen saturation) and clinical concepts (e.g., "weaning," "extubation") to event risk. Our variable-length multimodal framework improves performance and provides timestep-level feature importance, enhancing explainability and clinical relevance of deep learning models in critical care.

Medical subject headings