Extracting structured information from unstructured data is one of the key challenges in modern information retrieval applications, including e-commerce. Here, we demonstrate how recent advances in machine learning, combined with a recently published multilingual data set with standardized fine-grained product category information, enable robust product attribute extraction in challenging transfer learning settings. Our models can reliably predict product attributes across online shops, languages, or both. Furthermore, we show that our models can be used to match product taxonomies between online retailers.
翻译:从非结构化数据中提取结构化信息是现代信息检索应用(包括电子商务)中的关键挑战之一。我们在此展示了如何利用机器学习的最新进展,结合近期发布的、包含标准化细粒度产品类别信息的多语言数据集,在具有挑战性的迁移学习场景中实现稳健的产品属性提取。我们的模型能够可靠地跨在线商店、跨语言或同时跨两者预测产品属性。此外,我们还表明,这些模型可用于匹配在线零售商之间的产品分类体系。