Efficiently learning visual representations of items is vital for large-scale recommendations. In this article we compare several pretrained efficient backbone architectures, both in the convolutional neural network (CNN) and in the vision transformer (ViT) family. We describe challenges in e-commerce vision applications at scale and highlight methods to efficiently train, evaluate, and serve visual representations. We present ablation studies evaluating visual representations in several downstream tasks. To this end, we present a novel multilingual text-to-image generative offline evaluation method for visually similar recommendation systems. Finally, we include online results from deployed machine learning systems in production on a large scale e-commerce platform.
翻译:高效学习商品的视觉表征对于大规模推荐系统至关重要。本文对比了多种预训练的高效骨干架构,涵盖卷积神经网络与视觉Transformer家族。我们阐述了大规模电子商务视觉应用中的挑战,并重点介绍了高效训练、评估及部署视觉表征的方法。我们通过消融实验评估了视觉表征在多项下游任务中的表现。为此,我们提出了一种新颖的多语言文本到图像生成式离线评估方法,用于视觉相似推荐系统。最后,我们展示了在大型电子商务平台生产环境中部署的机器学习系统的在线实验结果。