Deep kernel learning enhanced fusion for multimodal classification.

Zhang, Duoyi; Bashar, Md Abul; Nayak, Richi · Neural Netw · 2026

basic_science · Level V

Where this comes from

Abstract

Multimodal learning aims to combine multiple modalities into a joint representation space. Existing works primarily focus on designing multimodal architectures to capture task-specific multimodal features. However, these methods often overlook variations in informativeness across modalities, leading to the feature-level bias problem, where irrelevant information introduces noise into the joint representation space. To address this issue, we propose a novel Deep Kernel Learning enhanced Multimodal Classification (DKLMC) framework. The framework comprises a main network for multimodal classification and a deep kernel learning network for estimating the unimodal contribution for each instance. The DKLMC framework is capable of estimating the contribution of each modality for each instance based on its unimodal contribution factor. This factor is derived using an auxiliary network, and each modality's informativeness is quantified by incorporating the uncertainty of the prediction confidence. Specifically, it addresses the feature-level bias problem by emphasising the more informative unimodal feature and reducing the impact of less informative features in the joint representation space. It features two key innovations: (1) an auxiliary network based on deep kernel learning to estimate unimodal contribution factor effectively, and (2) the use of the Reptile algorithm, a gradient-based meta-learning approach that aims to improve adaptability across various tasks, to balance the optimisation between main and auxiliary networks, ensuring effective learning of both unimodal contribution estimation and multimodal classification tasks. Extensive experiments across diverse datasets demonstrate the effectiveness of the proposed DKLMC framework.