As the content on the Internet continues to grow, many new dynamically changing and heterogeneous sources of data constantly emerge. A conventional search engine cannot crawl and index at the same pace as the expansion of the Internet. Moreover, a large portion of the data on the Internet is not accessible to traditional search engines. Distributed Information Retrieval (DIR) is a viable solution to this as it integrates multiple shards (resources) and provides a unified access to them. Resource selection is a key component of DIR systems. There is a rich body of literature on resource selection approaches for DIR. A key limitation of the existing approaches is that they primarily use term-based statistical features and do not generally model resource-query and resource-resource relationships. In this paper, we propose a graph neural network (GNN) based approach to learning-to-rank that is capable of modeling resource-query and resource-resource relationships. Specifically, we utilize a pre-trained language model (PTLM) to obtain semantic information from queries and resources. Then, we explicitly build a heterogeneous graph to preserve structural information of query-resource relationships and employ GNN to extract structural information. In addition, the heterogeneous graph is enriched with resource-resource type of edges to further enhance the ranking accuracy. Extensive experiments on benchmark datasets show that our proposed approach is highly effective in resource selection. Our method outperforms the state-of-the-art by 6.4% to 42% on various performance metrics.
翻译:随着互联网内容持续增长,大量动态变化且异构的数据源不断涌现。传统搜索引擎无法与互联网的扩展速度同步进行抓取和索引。此外,互联网上的大部分数据对传统搜索引擎不可见。分布式信息检索(DIR)是一种可行的解决方案,它整合多个分片(资源)并提供统一访问。资源选择是DIR系统的关键组件。现有文献中有大量关于DIR资源选择方法的研究。这些方法的一个关键局限性在于它们主要使用基于术语的统计特征,通常不建模资源-查询和资源-资源关系。本文提出了一种基于图神经网络(GNN)的学习排序方法,能够对资源-查询和资源-资源关系进行建模。具体而言,我们利用预训练语言模型(PTLM)获取查询和资源的语义信息,然后显式构建异构图以保留查询-资源关系的结构信息,并采用GNN提取结构特征。此外,异构图通过引入资源-资源类型的边进一步丰富,以提升排序精度。在基准数据集上的大量实验表明,我们的方法在资源选择上具有高效性。该方法在多种性能指标上以6.4%至42%的优势超越现有最优技术。