Information retrieval (IR) or knowledge retrieval, is a critical component for many down-stream tasks such as open-domain question answering (QA). It is also very challenging, as it requires succinctness, completeness, and correctness. In recent works, dense retrieval models have achieved state-of-the-art (SOTA) performance on in-domain IR and QA benchmarks by representing queries and knowledge passages with dense vectors and learning the lexical and semantic similarity. However, using single dense vectors and end-to-end supervision are not always optimal because queries may require attention to multiple aspects and event implicit knowledge. In this work, we propose an information retrieval pipeline that uses entity/event linking model and query decomposition model to focus more accurately on different information units of the query. We show that, while being more interpretable and reliable, our proposed pipeline significantly improves passage coverages and denotation accuracies across five IR and QA benchmarks. It will be the go-to system to use for applications that need to perform IR on a new domain without much dedicated effort, because of its superior interpretability and cross-domain performance.
翻译:信息检索或知识检索是许多下游任务(如开放域问答)的关键组成部分。这一任务极具挑战性,因为它要求检索结果具备简洁性、完整性和正确性。近年来的研究中,密集检索模型通过将查询和知识段落表示为密集向量,并学习词汇与语义相似性,已在域内信息检索和问答基准测试中取得最先进性能。然而,使用单一密集向量和端到端监督并非始终最优,因为查询可能需要关注多方面信息以及事件隐含知识。本文提出了一种信息检索流水线,该流水线结合实体/事件链接模型与查询分解模型,能够更精准地聚焦查询的不同信息单元。研究表明,我们提出的流水线在更具可解释性和可靠性的同时,在五个信息检索与问答基准测试中显著提升了段落覆盖范围和指称准确率。凭借其卓越的可解释性和跨领域性能,该流水线将成为需在新领域执行信息检索的任务中无需大量专门投入即可使用的首选系统。