Entity Set Expansion (ESE) aims to identify new entities belonging to the same semantic class as a given set of seed entities. Traditional methods primarily relied on positive seed entities to represent a target semantic class, which poses challenge for the representation of ultra-fine-grained semantic classes. Ultra-fine-grained semantic classes are defined based on fine-grained semantic classes with more specific attribute constraints. Describing it with positive seed entities alone cause two issues: (i) Ambiguity among ultra-fine-grained semantic classes. (ii) Inability to define "unwanted" semantic. Due to these inherent shortcomings, previous methods struggle to address the ultra-fine-grained ESE (Ultra-ESE). To solve this issue, we first introduce negative seed entities in the inputs, which belong to the same fine-grained semantic class as the positive seed entities but differ in certain attributes. Negative seed entities eliminate the semantic ambiguity by contrast between positive and negative attributes. Meanwhile, it provide a straightforward way to express "unwanted". To assess model performance in Ultra-ESE, we constructed UltraWiki, the first large-scale dataset tailored for Ultra-ESE. UltraWiki encompasses 236 ultra-fine-grained semantic classes, where each query of them is represented with 3-5 positive and negative seed entities. A retrieval-based framework RetExpan and a generation-based framework GenExpan are proposed to comprehensively assess the efficacy of large language models from two different paradigms in Ultra-ESE. Moreover, we devised three strategies to enhance models' comprehension of ultra-fine-grained entities semantics: contrastive learning, retrieval augmentation, and chain-of-thought reasoning. Extensive experiments confirm the effectiveness of our proposed strategies and also reveal that there remains a large space for improvement in Ultra-ESE.
翻译:摘要:实体集合扩展(ESE)旨在识别与给定种子实体集合属于同一语义类别的全新实体。传统方法主要依赖正种子实体来表征目标语义类别,这在超细粒度语义类别的表征中面临挑战。超细粒度语义类别是基于细粒度语义类别并附加更具体属性约束定义的。仅使用正种子实体进行描述会导致两个问题:(i)超细粒度语义类别间的歧义性;(ii)无法定义"非期望"语义。由于这些固有缺陷,现有方法难以解决超细粒度ESE(Ultra-ESE)问题。为解决该问题,我们首次在输入中引入负种子实体——这些实体与正种子实体属于同一细粒度语义类别,但在某些属性上存在差异。负种子实体通过正负属性的对比消除了语义歧义,同时提供了表达"非期望"语义的直观方式。为评估Ultra-ESE中模型性能,我们构建了UltraWiki——首个专为Ultra-ESE设计的大规模数据集。UltraWiki涵盖236个超细粒度语义类别,每个查询由3-5个正负种子实体共同表征。我们提出基于检索的框架RetExpan和基于生成的框架GenExpan,从两种不同范式全面评估大语言模型在Ultra-ESE中的效能。此外,我们设计了三种增强模型对超细粒度实体语义理解的策略:对比学习、检索增强与思维链推理。大量实验验证了所提策略的有效性,同时表明Ultra-ESE仍存在巨大的改进空间。