Entity matching, a core data integration problem, is the task of deciding whether two data tuples refer to the same real-world entity. Recent advances in deep learning methods, using pre-trained language models, were proposed for resolving entity matching. Although demonstrating unprecedented results, these solutions suffer from a major drawback as they require large amounts of labeled data for training, and, as such, are inadequate to be applied to low resource entity matching problems. To overcome the challenge of obtaining sufficient labeled data we offer a new active learning approach, focusing on a selection mechanism that exploits unique properties of entity matching. We argue that a distributed representation of a tuple pair indicates its informativeness when considered among other pairs. This is used consequently in our approach that iteratively utilizes space-aware considerations. Bringing it all together, we treat the low resource entity matching problem as a Battleship game, hunting indicative samples, focusing on positive ones, through awareness of the latent space along with careful planning of next sampling iterations. An extensive experimental analysis shows that the proposed algorithm outperforms state-of-the-art active learning solutions to low resource entity matching, and although using less samples, can be as successful as state-of-the-art fully trained known algorithms.
翻译:实体匹配是核心数据整合问题,其目标是判断两个数据元组是否指向同一现实世界实体。近年来,基于预训练语言模型的深度学习方法被提出用于解决实体匹配问题。尽管这些方法展现出前所未有的性能,但存在重大缺陷:它们需要大量标注数据进行训练,因此难以应用于低资源实体匹配问题。为克服标注数据不足的挑战,我们提出一种新型主动学习方法,其核心在于利用实体匹配独特属性的选择机制。我们认为,当将元组对置于其他对中考虑时,其分布式表示能指示信息量。据此,我们的方法迭代地采用空间感知策略。综合而言,我们将低资源实体匹配问题类比为战列舰游戏,通过潜空间感知与下一轮采样迭代的审慎规划,重点捕获指示性样本(尤其关注正样本)。大量实验分析表明,所提算法在低资源实体匹配任务上优于现有最先进主动学习方法,且在使用更少样本的情况下,能达到与完全训练的已知最先进算法相媲美的效果。