Complex logical query answering (CLQA) is a recently emerged task of graph machine learning that goes beyond simple one-hop link prediction and solves a far more complex task of multi-hop logical reasoning over massive, potentially incomplete graphs in a latent space. The task received a significant traction in the community; numerous works expanded the field along theoretical and practical axes to tackle different types of complex queries and graph modalities with efficient systems. In this paper, we provide a holistic survey of CLQA with a detailed taxonomy studying the field from multiple angles, including graph types (modality, reasoning domain, background semantics), modeling aspects (encoder, processor, decoder), supported queries (operators, patterns, projected variables), datasets, evaluation metrics, and applications. Refining the CLQA task, we introduce the concept of Neural Graph Databases (NGDBs). Extending the idea of graph databases (graph DBs), NGDB consists of a Neural Graph Storage and a Neural Graph Engine. Inside Neural Graph Storage, we design a graph store, a feature store, and further embed information in a latent embedding store using an encoder. Given a query, Neural Query Engine learns how to perform query planning and execution in order to efficiently retrieve the correct results by interacting with the Neural Graph Storage. Compared with traditional graph DBs, NGDBs allow for a flexible and unified modeling of features in diverse modalities using the embedding store. Moreover, when the graph is incomplete, they can provide robust retrieval of answers which a normal graph DB cannot recover. Finally, we point out promising directions, unsolved problems and applications of NGDB for future research.
翻译:复杂逻辑查询解答(CLQA)是图机器学习领域近期兴起的一项任务,它超越了简单的单跳链接预测,旨在解决在潜在空间中针对大规模、可能不完整的图进行多跳逻辑推理这一更为复杂的任务。该任务在学术界引起了广泛关注;众多研究沿理论和实践两个方向拓展了该领域,以高效系统处理不同类型复杂查询和图模态。本文对CLQA进行了全面综述,通过详细的分类从多角度审视该领域,包括图类型(模态、推理域、背景语义)、建模方面(编码器、处理器、解码器)、支持的查询(算子、模式、投影变量)、数据集、评估指标及应用。在细化CLQA任务的基础上,我们提出了神经图数据库(NGDBs)的概念。NGDB扩展了传统图数据库的思想,由神经图存储和神经图引擎组成。在神经图存储内部,我们设计了图存储、特征存储,并通过编码器将信息进一步嵌入潜在嵌入存储中。针对给定查询,神经查询引擎学习如何执行查询规划与执行,通过与神经图存储交互高效检索正确结果。与传统图数据库相比,NGDB利用嵌入存储实现了对不同模态特征的灵活统一建模。此外,当图不完整时,NGDB可提供传统图数据库无法恢复的鲁棒答案检索。最后,我们指出了NGDB未来研究中有前景的方向、未解决问题及应用场景。