This paper introduces ENEIDE (Extracting Named Entities from Italian Digital Editions), a silver standard dataset for Named Entity Recognition and Linking (NERL) in historical Italian texts. The corpus comprises 2,111 documents with over 8,000 entity annotations semi-automatically extracted from two scholarly digital editions: Digital Zibaldone, the philosophical diary of the Italian poet Giacomo Leopardi (1798--1837), and Aldo Moro Digitale, the complete works of the Italian politician Aldo Moro (1916--1978). Annotations cover multiple entity types (person, location, organization, literary work) linked to Wikidata identifiers, including NIL entities that cannot be mapped to the knowledge graph. To the best of our knowledge, ENEIDE represents the first multi-domain, publicly available NERL dataset for historical Italian with training, development, and test splits. We present a methodology for semi-automatic annotations extraction from manually curated scholarly digital editions, including quality control and annotation enhancement procedures. Baseline experiments using state-of-the-art models demonstrate the dataset's challenge for NERL and the gap between zero-shot approaches and fine-tuned models. The dataset's diachronic coverage spanning two centuries makes it particularly suitable for temporal entity disambiguation and cross-domain evaluation. ENEIDE is released under a CC BY-NC-SA 4.0 license.
翻译:本文介绍了ENEIDE(从意大利语数字版本中提取命名实体),一个用于历史意大利语文本中命名实体识别与链接(NERL)的银标准数据集。该语料库包含2,111篇文档,拥有超过8,000条实体标注,这些标注从两个学术数字版本中半自动化提取:意大利诗人贾科莫·莱奥帕尔迪(1798–1837)的哲学日记《数字齐巴尔迪》,以及意大利政治家阿尔多·莫罗(1916–1978)的全集《阿尔多·莫罗数字版》。标注涵盖多种实体类型(人物、地点、组织、文学作品),并与维基数据标识符关联,其中包括无法映射至知识图谱的NIL实体。据我们所知,ENEIDE是首个面向历史意大利语的多领域、公开可用的NERL数据集,提供了训练集、开发集和测试集划分。我们提出了一种从人工精心策划的学术数字版本中半自动化提取标注的方法论,包括质量控制和标注增强流程。使用最先进模型进行的基线实验展示了该数据集对NERL的挑战性,以及零样本方法与微调模型之间的性能差距。该数据集覆盖两个世纪的时间跨度,使其特别适用于时间实体消歧和跨领域评估。ENEIDE以CC BY-NC-SA 4.0许可证发布。