Answering complex queries on incomplete knowledge graphs is a challenging task where a model needs to answer complex logical queries in the presence of missing knowledge. Recently, Arakelyan et al. (2021); Minervini et al. (2022) showed that neural link predictors could also be used for answering complex queries: their Continuous Query Decomposition (CQD) method works by decomposing complex queries into atomic sub-queries, answers them using neural link predictors and aggregates their scores via t-norms for ranking the answers to each complex query. However, CQD does not handle negations and only uses the training signal from atomic training queries: neural link prediction scores are not calibrated to interact together via fuzzy logic t-norms during complex query answering. In this work, we propose to address this problem by training a parameter-efficient score adaptation model to re-calibrate neural link prediction scores: this new component is trained on complex queries by back-propagating through the complex query-answering process. Our method, CQD$^{A}$, produces significantly more accurate results than current state-of-the-art methods, improving from $34.4$ to $35.1$ Mean Reciprocal Rank values averaged across all datasets and query types while using $\leq 35\%$ of the available training query types. We further show that CQD$^{A}$ is data-efficient, achieving competitive results with only $1\%$ of the training data, and robust in out-of-domain evaluations.
翻译:在知识图谱不完整的情况下回答复杂查询是一项具有挑战性的任务,模型需要在存在知识缺失的条件下处理复杂逻辑查询。近期,Arakelyan等人(2021)和Minervini等人(2022)的研究表明,神经链接预测器也可用于回答复杂查询:其提出的连续查询分解(CQD)方法通过将复杂查询分解为原子子查询,利用神经链接预测器进行回答,并采用t-范数聚合各子查询得分以对每个复杂查询的答案进行排序。然而,CQD方法无法处理否定操作,且仅利用原子训练查询的监督信号:在复杂查询回答过程中,神经链接预测得分未经过校准以通过模糊逻辑t-范数实现协同作用。针对这一问题,本文提出通过训练一个参数高效的得分自适应模型来重新校准神经链接预测得分:该新组件通过反向传播复杂查询回答过程,在复杂查询上进行训练。我们的方法CQD$^{A}$在所有数据集和查询类型上的平均倒数排名从$34.4$提升至$35.1$,同时仅使用$\leq 35\%$的可用训练查询类型,其准确率显著优于当前最先进方法。我们进一步证明CQD$^{A}$具有数据高效性——仅需$1\%$的训练数据即可取得具有竞争力的结果,并在域外评估中展现出鲁棒性。