Entity alignment (EA) aims at identifying equivalent entity pairs across different knowledge graphs (KGs) that refer to the same real-world identity. To systematically combat confirmation bias for pseudo-labeling-based entity alignment, we propose a Unified Pseudo-Labeling framework for Entity Alignment (UPL-EA) that explicitly eliminates pseudo-labeling errors to boost the accuracy of entity alignment. UPL-EA consists of two complementary components: (1) The Optimal Transport (OT)-based pseudo-labeling uses discrete OT modeling as an effective means to enable more accurate determination of entity correspondences across two KGs and to mitigate the adverse impact of erroneous matches. A simple but highly effective criterion is further devised to derive pseudo-labeled entity pairs that satisfy one-to-one correspondences at each iteration. (2) The cross-iteration pseudo-label calibration operates across multiple consecutive iterations to further improve the pseudo-labeling precision rate by reducing the local pseudo-label selection variability with a theoretical guarantee. The two components are respectively designed to eliminate Type I and Type II pseudo-labeling errors identified through our analyse. The calibrated pseudo-labels are thereafter used to augment prior alignment seeds to reinforce subsequent model training for alignment inference. The effectiveness of UPL-EA in eliminating pseudo-labeling errors is both theoretically supported and experimentally validated. The experimental results show that our approach achieves competitive performance with limited prior alignment seeds.
翻译:实体对齐旨在识别不同知识图谱中指向同一真实世界实体的等价实体对。为系统性地消除基于伪标记的实体对齐中存在的确认偏差,我们提出了一种面向实体对齐的统一伪标记框架,该框架通过显式消除伪标记错误来提升实体对齐的准确性。该框架包含两个互补组件:(1)基于最优传输的伪标记方法利用离散最优传输建模作为有效手段,更精确地确定两个知识图谱间的实体对应关系,并减轻错误匹配的不利影响。我们进一步设计了一个简单但高效的准则,用于在每次迭代中推导满足一一对应关系的伪标记实体对。(2)跨迭代伪标记校准通过跨多个连续迭代运行,在理论上保证降低局部伪标记选择变异性的前提下,进一步提升伪标记精度。这两个组件分别针对我们分析中识别的第一类与第二类伪标记消除错误。校准后的伪标记随后用于增强先前的对齐种子,以强化后续用于对齐推理的模型训练。本文从理论上支持并实验验证了所提框架消除伪标记错误的有效性。实验结果表明,在有限的先验对齐种子条件下,本方法取得了具有竞争力的性能。