Given the impact of language models on the field of Natural Language Processing, a number of Spanish encoder-only masked language models (aka BERTs) have been trained and released. These models were developed either within large projects using very large private corpora or by means of smaller scale academic efforts leveraging freely available data. In this paper we present a comprehensive head-to-head comparison of language models for Spanish with the following results: (i) Previously ignored multilingual models from large companies fare better than monolingual models, substantially changing the evaluation landscape of language models in Spanish; (ii) Results across the monolingual models are not conclusive, with supposedly smaller and inferior models performing competitively. Based on these empirical results, we argue for the need of more research to understand the factors underlying them. In this sense, the effect of corpus size, quality and pre-training techniques need to be further investigated to be able to obtain Spanish monolingual models significantly better than the multilingual ones released by large private companies, specially in the face of rapid ongoing progress in the field. The recent activity in the development of language technology for Spanish is to be welcomed, but our results show that building language models remains an open, resource-heavy problem which requires to marry resources (monetary and/or computational) with the best research expertise and practice.
翻译:鉴于语言模型对自然语言处理领域的影响,目前已训练并发布了多个西班牙语编码器仅掩码语言模型(即BERT类模型)。这些模型的开发路径包括两类:一类依托大型项目使用超大规模私有语料库,另一类则通过学术团队利用公开可用数据开展小型研究。本文对西班牙语语言模型进行了系统性横向对比,获得以下发现:(i) 大型企业此前被忽视的多语言模型表现优于单语言模型,显著改变了西班牙语语言模型的评估格局;(ii) 单语言模型间的对比结果缺乏确定性,预期规模较小且性能较弱的模型反而展现出竞争力。基于这些实证结果,我们主张需进一步研究以解析其背后成因。具体而言,语料库规模、质量及预训练技术的影响仍需深入探究,方能构建显著优于大型私企发布的多语言模型的西班牙语单语言模型——尤其在当前领域快速发展的背景下。近期西班牙语语言技术研发活动值得肯定,但研究结果表明:构建语言模型仍是一项资源密集型开放性难题,需要将资金与计算资源同顶尖研究能力及实践经验相结合。