Live commerce is the act of selling products online through live streaming. The customer's diverse demands for online products introduce more challenges to Livestreaming Product Recognition. Previous works have primarily focused on fashion clothing data or utilize single-modal input, which does not reflect the real-world scenario where multimodal data from various categories are present. In this paper, we present LPR4M, a large-scale multimodal dataset that covers 34 categories, comprises 3 modalities (image, video, and text), and is 50? larger than the largest publicly available dataset. LPR4M contains diverse videos and noise modality pairs while exhibiting a long-tailed distribution, resembling real-world problems. Moreover, a cRoss-vIew semantiC alignmEnt (RICE) model is proposed to learn discriminative instance features from the image and video views of the products. This is achieved through instance-level contrastive learning and cross-view patch-level feature propagation. A novel Patch Feature Reconstruction loss is proposed to penalize the semantic misalignment between cross-view patches. Extensive experiments demonstrate the effectiveness of RICE and provide insights into the importance of dataset diversity and expressivity. The dataset and code are available at https://github.com/adxcreative/RICE
翻译:直播电商是通过直播平台在线销售商品的行为。顾客对线上商品的多样化需求为直播商品识别带来了更多挑战。现有工作主要聚焦于时尚服饰数据或采用单模态输入,未能反映现实场景中多类别多模态数据并存的情况。本文提出大规模多模态数据集LPR4M,涵盖34个类别、包含图像、视频和文本三种模态,其规模比现有最大公开数据集大50%以上。LPR4M包含多样化的视频与噪声模态配对,同时呈现长尾分布特性,贴近现实问题。此外,提出跨视图语义对齐模型RICE,通过学习商品图像与视频视图的判别性实例特征,采用实例级对比学习与跨视图分块级特征传播实现该目标。提出新型分块特征重构损失函数,用于惩罚跨视图分块间的语义错位。大量实验验证了RICE的有效性,并揭示了数据集多样性与表现力的重要性。数据集与代码已开源至https://github.com/adxcreative/RICE