To advance the neural encoding of Portuguese (PT), and a fortiori the technological preparation of this language for the digital age, we developed a Transformer-based foundation model that sets a new state of the art in this respect for two of its variants, namely European Portuguese from Portugal (PT-PT) and American Portuguese from Brazil (PT-BR). To develop this encoder, which we named Albertina PT-*, a strong model was used as a starting point, DeBERTa, and its pre-training was done over data sets of Portuguese, namely over a data set we gathered for PT-PT and over the brWaC corpus for PT-BR. The performance of Albertina and competing models was assessed by evaluating them on prominent downstream language processing tasks adapted for Portuguese. Both Albertina PT-PT and PT-BR versions are distributed free of charge and under the most permissive license possible and can be run on consumer-grade hardware, thus seeking to contribute to the advancement of research and innovation in language technology for Portuguese.
翻译:为推进葡萄牙语的神经编码,进而提升该语言在数字时代的技术准备水平,我们开发了一个基于Transformer的基础模型,在葡萄牙语的两种变体——即葡萄牙欧洲葡萄牙语(PT-PT)和巴西美洲葡萄牙语(PT-BR)方面,该模型设立了新的最优水平。为开发这一编码器(我们命名为Albertina PT-*),我们以强模型DeBERTa为起点,并使用葡萄牙语数据集进行预训练,具体包括我们为PT-PT收集的数据集以及PT-BR的brWaC语料库。通过评估Albertina及对比模型在针对葡萄牙语适配的突出下游语言处理任务上的表现,我们对其性能进行了评测。Albertina PT-PT和PT-BR两个版本均以最宽松的许可免费分发,并可在消费级硬件上运行,旨在为葡萄牙语语言技术的研究与创新做出贡献。