Query Performance Prediction (QPP) estimates the effectiveness of a search engine's results in response to a query without relevance judgments. Traditionally, post-retrieval predictors have focused upon either the distribution of the retrieval scores, or the coherence of the top-ranked documents using traditional bag-of-words index representations. More recently, BERT-based models using dense embedded document representations have been used to create new predictors, but mostly applied to predict the performance of rankings created by BM25. Instead, we aim to predict the effectiveness of rankings created by single-representation dense retrieval models (ANCE & TCT-ColBERT). Therefore, we propose a number of variants of existing unsupervised coherence-based predictors that employ neural embedding representations. In our experiments on the TREC Deep Learning Track datasets, we demonstrate improved accuracy upon dense retrieval (up to 92% compared to sparse variants for TCT-ColBERT and 188% for ANCE). Going deeper, we select the most representative and best performing predictors to study the importance of differences among predictors and query types on query performance. Using existing distribution-based evaluation QPP measures and a particular type of linear mixed models, we find that query types further significantly influence query performance (and are up to 35% responsible for the unstable performance of QPP predictors), and that this sensitivity is unique to dense retrieval models. Our approach introduces a new setting for obtaining richer information on query differences in dense QPP that can explain potential unstable performance of existing predictors and outlines the unique characteristics of different query types on dense retrieval models.
翻译:查询性能预测(QPP)可在无需相关性判断的情况下估计搜索引擎针对某查询所返回结果的有效性。传统上,后检索预测器主要关注检索得分的分布,或基于传统词袋索引表示的前置文档一致性。近年来,基于BERT的稠密嵌入文档表示模型已被用于构建新型预测器,但其主要应用于预测由BM25生成排名的性能。本研究旨在预测由单表示稠密检索模型(ANCE与TCT-ColBERT)生成排名的有效性。为此,我们提出了一系列基于神经嵌入表示的现有无监督一致性预测器变体。在TREC深度学习赛道数据集上的实验表明,本方法在稠密检索上显著提升了预测精度(与稀疏变体相比,TCT-ColBERT提升达92%,ANCE提升达188%)。进一步地,我们选取最具代表性且性能最优的预测器,以探究预测器差异及查询类型对查询性能的影响。通过结合现有基于分布的QPP评估指标与特定线性混合模型,我们发现查询类型会显著影响查询性能(其对QPP预测器不稳定的性能贡献度高达35%),且这种敏感性仅存在于稠密检索模型中。本方法为稠密QPP中获取更丰富的查询差异信息提供了新范式,可解释现有预测器潜在的不稳定性能,并揭示了不同查询类型在稠密检索模型中的独特特征。