Matching identical products present in multiple product feeds constitutes a crucial element of many tasks of e-commerce, such as comparing product offerings, dynamic price optimization, and selecting the assortment personalized for the client. It corresponds to the well-known machine learning task of entity matching, with its own specificity, like omnipresent unstructured data or inaccurate and inconsistent product descriptions. This paper aims to present a new philosophy to product matching utilizing a semi-supervised clustering approach. We study the properties of this method by experimenting with the IDEC algorithm on the real-world dataset using predominantly textual features and fuzzy string matching, with more standard approaches as a point of reference. Encouraging results show that unsupervised matching, enriched with a small annotated sample of product links, could be a possible alternative to the dominant supervised strategy, requiring extensive manual data labeling.
翻译:匹配多个产品源中的相同产品是电商诸多任务的关键环节,例如产品对比、动态价格优化以及为客户个性化筛选商品组合。这一任务对应机器学习领域经典的实体匹配问题,但具有其特殊性,例如普遍存在的非结构化数据、不准确且不一致的产品描述。本文旨在提出一种利用半监督聚类方法进行产品匹配的新思路。我们通过在真实数据集上使用IDEC算法进行实验,主要采用文本特征和模糊字符串匹配,并以更标准的方法作为参照,研究了该方法的特性。令人鼓舞的结果表明,通过少量标注产品链接进行增强的无监督匹配,有可能成为需要大量人工数据标注的主流监督策略的替代方案。