We present the development of a Named Entity Recognition (NER) dataset for Tagalog. This corpus helps fill the resource gap present in Philippine languages today, where NER resources are scarce. The texts were obtained from a pretraining corpora containing news reports, and were labeled by native speakers in an iterative fashion. The resulting dataset contains ~7.8k documents across three entity types: Person, Organization, and Location. The inter-annotator agreement, as measured by Cohen's $\kappa$, is 0.81. We also conducted extensive empirical evaluation of state-of-the-art methods across supervised and transfer learning settings. Finally, we released the data and processing code publicly to inspire future work on Tagalog NLP.
翻译:我们报告了针对他加禄语命名实体识别(NER)数据集的开发工作。该语料库有助于填补当前菲律宾语言中NER资源匮乏的空白。文本数据来源于包含新闻报道的预训练语料库,并由母语者以迭代方式完成标注。最终数据集包含约7800篇文档,涵盖三类实体:人物、组织和地点。通过Cohen's κ系数衡量的标注者间一致性达到0.81。我们还对监督学习和迁移学习场景下的前沿方法进行了广泛的实证评估。最后,我们将数据和处理代码公开发布,以推动他加禄语自然语言处理领域的后续研究。