Knowledge graph representation learning (KGRL) or knowledge graph embedding (KGE) plays a crucial role in AI applications for knowledge construction and information exploration. These models aim to encode entities and relations present in a knowledge graph into a lower-dimensional vector space. During the training process of KGE models, using positive and negative samples becomes essential for discrimination purposes. However, obtaining negative samples directly from existing knowledge graphs poses a challenge, emphasizing the need for effective generation techniques. The quality of these negative samples greatly impacts the accuracy of the learned embeddings, making their generation a critical aspect of KGRL. This comprehensive survey paper systematically reviews various negative sampling (NS) methods and their contributions to the success of KGRL. Their respective advantages and disadvantages are outlined by categorizing existing NS methods into five distinct categories. Moreover, this survey identifies open research questions that serve as potential directions for future investigations. By offering a generalization and alignment of fundamental NS concepts, this survey provides valuable insights for designing effective NS methods in the context of KGRL and serves as a motivating force for further advancements in the field.
翻译:知识图谱表示学习或知识图谱嵌入在知识构建和信息探索的人工智能应用中扮演着关键角色。这些模型旨在将知识图谱中的实体和关系编码到低维向量空间中。在知识图谱嵌入模型的训练过程中,使用正样本和负样本对于区分任务至关重要。然而,直接从现有知识图谱获取负样本存在挑战,凸显了有效生成技术的必要性。负样本的质量显著影响学习嵌入的准确性,使其生成成为知识图谱表示学习的关键环节。本篇综合综述论文系统地回顾了各类负采样方法及其对知识图谱表示学习成功所做的贡献。通过将现有负采样方法划分为五个不同类别,本文概述了它们各自的优缺点。此外,本综述指出了可作为未来研究潜在方向的开放性问题。通过对负采样基本概念的泛化与对齐,本综述为在知识图谱表示学习背景下设计有效的负采样方法提供了宝贵见解,并推动了该领域的进一步发展。