Due to the massive size of test collections, a standard practice in IR evaluation is to construct a 'pool' of candidate relevant documents comprised of the top-k documents retrieved by a wide range of different retrieval systems - a process called depth-k pooling. A standard practice is to set the depth (k) to a constant value for each query constituting the benchmark set. However, in this paper we argue that the annotation effort can be substantially reduced if the depth of the pool is made a variable quantity for each query, the rationale being that the number of documents relevant to the information need can widely vary across queries. Our hypothesis is that a lower depth for the former class of queries and a higher depth for the latter can potentially reduce the annotation effort without a significant change in retrieval effectiveness evaluation. We make use of standard query performance prediction (QPP) techniques to estimate the number of potentially relevant documents for each query, which is then used to determine the depth of the pool. Our experiments conducted on standard test collections demonstrate that this proposed method of employing query-specific variable depths is able to adequately reflect the relative effectiveness of IR systems with a substantially smaller annotation effort.
翻译:由于测试集规模庞大,信息检索评估中的标准做法是构建一个由不同检索系统返回的前k篇文档组成的候选相关文档"池"——这一过程称为深度k池化。通常做法是将基准测试集中每个查询的深度(k)设置为恒定值。然而,本文提出若将池深度设为每个查询的可变量,则可大幅减少标注工作量,其依据是不同查询对应的信息需求相关文档数量可能存在巨大差异。我们的假设是:对前一类查询采用较低深度、后一类查询采用较高深度,可在不显著改变检索效果评估的前提下减少标注工作量。我们利用标准查询性能预测(QPP)技术来估算每个查询的潜在相关文档数量,并据此确定池深度。在标准测试集上开展的实验表明,这种采用查询特异可变深度的新方法能够以大幅减少的标注工作量充分反映信息检索系统的相对有效性。