ID-Guided Multimodal experts with contrastive diffusion for sequential recommendation.

Lu, Yi-Hong; Xi, Wu-Dong; Xing, Xing-Xing; Wan, Wei; Wang, Chang-Dong · Neural Netw · 2026

Where this comes from

Abstract

Multimodal sequential recommendation enriches user-item interaction modeling by incorporating text and image modalities. However, existing methods often overlook the inherent inconsistencies between different modalities and fail to effectively filter redundant noise within modality-specific features, leading to suboptimal recommendation performance. To address these issues, we propose a novel framework named ID-Guided Multimodal Experts with Contrastive Diffusion for Sequential Recommendation (IMECD). Specifically, IMECD introduces a novel ID-guided multimodal mixture of experts module, which uniquely leverages long-term user preferences encoded in ID embeddings to dynamically guide the extraction of text and image features. This module helps resolve cross-modal semantic inconsistency and suppresses irrelevant signals, thereby improving the quality of multimodal representations. To further mitigate noise in user interaction sequences, we introduce a modality-specific vector quantization module that denoises sequential features by independently quantizing each modality. Moreover, we propose a contrastive diffusion generation module, which conditions the diffusion process on sequence representations and employs a contrastive loss to alleviate generation bias. Extensive experiments on four benchmark datasets demonstrate that IMECD consistently outperforms state-of-the-art baselines. Our code is available at https://anonymous.4open.science/r/IMECD-LYH.

Medical subject headings