Estimating position bias is a well-known challenge in Learning to Rank (L2R). Click data in e-commerce applications, such as targeted advertisements and search engines, provides implicit but abundant feedback to improve personalized rankings. However, click data inherently includes various biases like position bias. Based on the position-based click model, Result Randomization and Regression Expectation-Maximization algorithm (REM) have been proposed to estimate position bias, but they require various paired observations of (item, position). In real-world scenarios of advertising, marketers frequently display advertisements in a fixed pre-determined order, which creates difficulties in estimation due to the limited availability of various pairs in the training data, resulting in a sparse dataset. We propose a variant of the REM that utilizes item embeddings to alleviate the sparsity of (item, position). Using a public dataset and internal carousel advertisement click dataset, we empirically show that item embedding with Latent Semantic Indexing (LSI) and Variational Auto-Encoder (VAE) improves the accuracy of position bias estimation and the estimated position bias enhances Learning to Rank performance. We also show that LSI is more effective as an embedding creation method for position bias estimation.
翻译:在排序学习(L2R)中,估计立场偏差是一个众所周知的挑战。电子商务应用中的点击数据(如定向广告和搜索引擎)提供了隐式但丰富的反馈,用于改进个性化排序。然而,点击数据本身包含各种偏差,例如立场偏差。基于基于立场的点击模型,已提出结果随机化和回归期望最大化算法(REM)来估计立场偏差,但该方法需要多种(物品,立场)配对观测。在实际广告场景中,营销人员经常以固定的预定顺序展示广告,由于训练数据中可用配对有限,导致数据稀疏,这给估计带来了困难。我们提出了一种REM的变体,利用物品嵌入来缓解(物品,立场)的稀疏性。通过使用公开数据集和内部轮播广告点击数据集,我们经验性地证明,使用潜在语义索引(LSI)和变分自编码器(VAE)的物品嵌入提高了立场偏差估计的准确性,并且估计出的立场偏差提升了排序学习性能。我们还表明,LSI作为一种用于立场偏差估计的嵌入创建方法更为有效。