The purpose of modeling document relevance for search engines is to rank better in subsequent searches. Document-specific historical click-through rates can be important features in a dynamic ranking system which updates as we accumulate more sample. This paper describes the properties of several such features, and tests them in controlled experiments. Extending the inverse propensity weighting method to documents creates an unbiased estimate of document relevance. This feature can approximate relevance accurately, leading to near-optimal ranking in ideal circumstances. However, it has high variance that is increasing with respect to the degree of position bias. Furthermore, inaccurate position bias estimation leads to poor performance. Under several scenarios this feature can perform worse than biased click-through rates. This paper underscores the need for accurate position bias estimation, and is unique in suggesting simultaneous use of biased and unbiased position bias features.
翻译:搜索引擎文档相关性建模的目标是优化后续搜索的排序效果。文档特定的历史点击率作为动态排序系统中的重要特征,会随样本积累而持续更新。本文描述了此类特征的若干特性,并通过受控实验进行验证。将逆倾向加权方法扩展至文档层面,可构建文档相关性的无偏估计量。该特征在理想条件下能精确近似相关性,实现接近最优的排序效果,但其方差较高且随位置偏差程度递增。此外,不准确的位置偏差估计将导致性能下降。在多种场景下,该特征的表现甚至劣于有偏点击率。本文强调了准确估计位置偏差的必要性,并首次提出联合使用有偏特征与无偏位置偏差特征的建议。