With the advancement of multimedia internet, the impact of visual characteristics on the decision of users to click or not within the online retail industry is increasingly significant. Thus, incorporating visual features is a promising direction for further performance improvements in click-through rate (CTR). However, experiments on our production system revealed that simply injecting the image embeddings trained with established pre-training methods only has marginal improvements. We believe that the main advantage of existing image feature pre-training methods lies in their effectiveness for cross-modal predictions. However, this differs significantly from the task of CTR prediction in recommendation systems. In recommendation systems, other modalities of information (such as text) can be directly used as features in downstream models. Even if the performance of cross-modal prediction tasks is excellent, it is challenging to provide significant information gain for the downstream models. We argue that a visual feature pre-training method tailored for recommendation is necessary for further improvements beyond existing modality features. To this end, we propose an effective user intention reconstruction module to mine visual features related to user interests from behavior histories, which constructs a many-to-one correspondence. We further propose a contrastive training method to learn the user intentions and prevent the collapse of embedding vectors. We conduct extensive experimental evaluations on public datasets and our production system to verify that our method can learn users' visual interests. Our method achieves $0.46\%$ improvement in offline AUC and $0.88\%$ improvement in Taobao GMV (Cross Merchandise Volume) with p-value$<$0.01.
翻译:随着多媒体互联网的发展,视觉特征对在线零售行业中用户点击决策的影响日益显著。因此,融合视觉特征成为进一步提升点击率(CTR)性能的重要方向。然而,在我们的生产系统实验中,直接注入基于现有预训练方法生成的图像嵌入仅能带来边际提升。我们认为,现有图像特征预训练方法的核心优势在于跨模态预测的有效性,但这与推荐系统中的CTR预测任务存在显著差异。在推荐系统中,其他模态信息(如文本)可直接作为下游模型的特征输入。即便跨模态预测任务表现优异,也难以向下游模型提供显著的信息增益。我们主张,针对推荐场景定制的视觉特征预训练方法才是突破现有模态特征瓶颈的关键。为此,我们提出了一种有效的用户意图重建模块,从行为历史中挖掘与用户兴趣相关的视觉特征,构建了多对一的对应关系。在此基础上,我们进一步提出对比训练方法以学习用户意图并防止嵌入向量坍缩。通过在公开数据集和生产系统上的大量实验验证,本方法能够有效学习用户的视觉兴趣。我们的方法在离线AUC上提升0.46%,在淘宝GMV(交叉商品成交额)上提升0.88%,且p值<0.01。