MMDOK: A Multi-modal and Multi-scale Disease-oriented Fusion Framework with Kolmogorov-Arnold Networks.

Li, Wei; Gong, Xun; Fan, Lin; Li, Xinxin; Li, Jiao; Sun, Xiaobin · IEEE J Biomed Health Inform · 2026

basic_science · Level V

Where this comes from

Abstract

Learning medical visual representations directly from image-report pairs has become an emerging topic in representation learning. However, the heterogeneity and complementary nature of medical reports and images pose challenges to adaptively fusing. We propose a Multi-modal and Multi-scale Disease-oriented fusion framework with Kolmogorov-Arnold Networks (MMDOK). This method utilizes the multi-scale features naturally present in medical images and reports and explores feature fusion between multi-view or multi-modal images and textual reports. Specifically, MMDOK designs a text encoder and two image encoders to enhance the diagnostic model by adaptively integrating text supervision and additional image information. To enhance the model's ability to extract information across different semantic scales, we introduce a Bidirectional Cross-Attention (BCA) module and a Cross-Modal Clustering (CMC) module. The proposed CMC module effectively captures disease consistent latent features from both images and clinical reports, alleviating overfitting and improving generalization to out-of-distribution test sets. To reduce cross-modal redundancy, we introduce a Disease-Oriented Attention (DOA) module, which adaptively assigns modality-specific weights based on disease labels, enabling more effective and context-aware feature fusion. Additionally, we embed Kolmogorov-Arnold Network layers into various modules to enable efficient and accurate feature extraction through enhanced nonlinear representation capacity. We evaluate our model's performance, achieving an average accuracy of 94.27% and 83.67% with and without text supervision on a public lung disease dataset, and 87.63% and 82.1% on a private submucosal tumors dataset, outperforming state-of-the-art alternatives.