Unbiased learning-to-rank (ULTR) is a well-established framework for learning from user clicks, which are often biased by the ranker collecting the data. While theoretically justified and extensively tested in simulation, ULTR techniques lack empirical validation, especially on modern search engines. The Baidu-ULTR dataset released for the WSDM Cup 2023, collected from Baidu's search engine, offers a rare opportunity to assess the real-world performance of prominent ULTR techniques. Despite multiple submissions during the WSDM Cup 2023 and the subsequent NTCIR ULTRE-2 task, it remains unclear whether the observed improvements stem from applying ULTR or other learning techniques. In this work, we revisit and extend the available experiments on the Baidu-ULTR dataset. We find that standard unbiased learning-to-rank techniques robustly improve click predictions but struggle to consistently improve ranking performance, especially considering the stark differences obtained by choice of ranking loss and query-document features. Our experiments reveal that gains in click prediction do not necessarily translate to enhanced ranking performance on expert relevance annotations, implying that conclusions strongly depend on how success is measured in this benchmark.
翻译:无偏学习排序(ULTR)是一个从用户点击中学习的成熟框架,而用户点击通常因收集数据的排序器而产生偏差。尽管在理论上得到验证并通过模拟进行了广泛测试,但ULTR技术缺乏经验验证,尤其是在现代搜索引擎上。为WSDM 2023杯发布的百度ULTR数据集,从百度搜索引擎中收集,为评估主流ULTR技术的真实世界性能提供了难得的机遇。尽管在WSDM 2023杯及随后的NTCIR ULTRE-2任务中提交了大量作品,但尚不清楚观察到的改进是源于应用ULTR还是其他学习技术。本研究重新审视并扩展了在百度ULTR数据集上已有的实验。我们发现,标准无偏学习排序技术能显著提升点击预测的鲁棒性,但在持续改善排序性能方面面临挑战,尤其是考虑到排序损失函数和查询-文档特征选择的显著差异。实验表明,点击预测的提升并不一定转化为基于专家相关性标注的排序性能增强,这意味着结论很大程度上取决于该基准测试中成功指标的衡量方式。