Language resources such as wordnets remain indispensable tools for different natural language tasks and applications. However, for low-resource languages such as Filipino, existing wordnets are old and outdated, and producing new ones may be slow and costly in terms of time and resources. In this paper, we propose an automatic method for constructing a wordnet from scratch using only an unlabeled corpus and a sentence embeddings-based language model. Using this, we produce FilWordNet, a new wordnet that supplants and improves the outdated Filipino WordNet. We evaluate our automatically-induced senses and synsets by matching them with senses from the Princeton WordNet, as well as comparing the synsets to the old Filipino WordNet. We empirically show that our method can induce existing, as well as potentially new, senses and synsets automatically without the need for human supervision.
翻译:诸如词网之类的语言资源仍然是不同自然语言任务和应用中不可或缺的工具。然而,对于菲律宾语等低资源语言,现有的词网老旧过时,而构建新词网在时间和资源方面可能既缓慢又昂贵。本文提出了一种仅使用未标注语料库和基于句子嵌入的语言模型,从零开始自动构建词网的方法。利用该方法,我们生成了FilWordNet——一个替代并改进陈旧菲律宾语词网的新词网。我们通过将自动归纳的词义与普林斯顿词网中的词义进行匹配,并将同义词集与旧版菲律宾语词网进行比较,从而评估了自动归纳的词义和同义词集。实验证明,我们的方法能够在无需人工监督的情况下,自动归纳出已有的以及潜在的新词义和同义词集。