Large multilingual models have inspired a new class of word alignment methods, which work well for the model's pretraining languages. However, the languages most in need of automatic alignment are low-resource and, thus, not typically included in the pretraining data. In this work, we ask: How do modern aligners perform on unseen languages, and are they better than traditional methods? We contribute gold-standard alignments for Bribri--Spanish, Guarani--Spanish, Quechua--Spanish, and Shipibo-Konibo--Spanish. With these, we evaluate state-of-the-art aligners with and without model adaptation to the target language. Finally, we also evaluate the resulting alignments extrinsically through two downstream tasks: named entity recognition and part-of-speech tagging. We find that although transformer-based methods generally outperform traditional models, the two classes of approach remain competitive with each other.
翻译:大型多语言模型催生了一类新的词对齐方法,这些方法在模型的预训练语言上表现良好。然而,最需要自动对齐的语言多为低资源语言,因此通常未被包含在预训练数据中。本研究提出疑问:现代对齐器在未见过的语言上表现如何?它们是否优于传统方法?我们为布里布里语-西班牙语、瓜拉尼语-西班牙语、克丘亚语-西班牙语和希皮博-科尼博语-西班牙语贡献了黄金标准对齐数据。基于这些数据,我们评估了最先进的对齐器(含/不含对目标语言的模型适配)。此外,我们还通过两个下游任务(命名实体识别和词性标注)对生成的对齐进行外部评估。研究发现,尽管基于Transformer的方法通常优于传统模型,但这两类方法仍保持相互竞争关系。