Sequential recommendation systems utilize the sequential interactions of users with items as their main supervision signals in learning users' preferences. However, existing methods usually generate unsatisfactory results due to the sparsity of user behavior data. To address this issue, we propose a novel pre-training framework, named Multimodal Sequence Mixup for Sequential Recommendation (MSM4SR), which leverages both users' sequential behaviors and items' multimodal content (\ie text and images) for effectively recommendation. Specifically, MSM4SR tokenizes each item image into multiple textual keywords and uses the pre-trained BERT model to obtain initial textual and visual features of items, for eliminating the discrepancy between the text and image modalities. A novel backbone network, \ie Multimodal Mixup Sequence Encoder (M$^2$SE), is proposed to bridge the gap between the item multimodal content and the user behavior, using a complementary sequence mixup strategy. In addition, two contrastive learning tasks are developed to assist M$^2$SE in learning generalized multimodal representations of the user behavior sequence. Extensive experiments on real-world datasets demonstrate that MSM4SR outperforms state-of-the-art recommendation methods. Moreover, we further verify the effectiveness of MSM4SR on other challenging tasks including cold-start and cross-domain recommendation.
翻译:序列推荐系统利用用户与物品的序列交互作为学习用户偏好的主要监督信号。然而,现有方法通常因用户行为数据的稀疏性而产生不理想的结果。为解决该问题,我们提出一种新颖的预训练框架——多模态序列混合用于序列推荐(MSM4SR),该框架同时利用用户的序列行为与物品的多模态内容(即文本和图像)进行有效推荐。具体而言,MSM4SR将每个物品图像切分为多个文本关键词,并采用预训练的BERT模型获取物品的初始文本与视觉特征,以消除文本与图像模态之间的差异。我们提出一种新型骨干网络——多模态混合序列编码器(M²SE),通过互补序列混合策略弥合物品多模态内容与用户行为之间的鸿沟。此外,我们开发了两个对比学习任务以辅助M²SE学习用户行为序列的泛化多模态表示。在真实世界数据集上的大量实验表明,MSM4SR优于最先进的推荐方法。同时,我们进一步验证了MSM4SR在冷启动与跨域推荐等其他挑战性任务上的有效性。