To foster the neural encoding of Portuguese, this paper contributes foundation encoder models that represent an expansion of the still very scarce ecosystem of large language models specifically developed for this language that are fully open, in the sense that they are open source and openly distributed for free under an open license for any purpose, thus including research and commercial usages. Like most languages other than English, Portuguese is low-resourced in terms of these foundational language resources, there being the inaugural 900 million parameter Albertina and 335 million Bertimbau. Taking this couple of models as an inaugural set, we present the extension of the ecosystem of state-of-the-art open encoders for Portuguese with a larger, top performance-driven model with 1.5 billion parameters, and a smaller, efficiency-driven model with 100 million parameters. While achieving this primary goal, further results that are relevant for this ecosystem were obtained as well, namely new datasets for Portuguese based on the SuperGLUE benchmark, which we also distribute openly.
翻译:为促进葡萄牙语的神经编码,本文贡献了基础编码器模型,这些模型扩展了专门为此语言开发的、完全开放的大型语言模型生态系统——目前该生态系统仍非常稀缺。所谓完全开放,指模型开源、免费公开分发、采用开放许可且可用于任何目的(包括研究和商业用途)。与英语之外的大多数语言类似,葡萄牙语在这些基础语言资源方面属于低资源语言,现有开创性模型包括9亿参数的Albertina和3.35亿参数的Bertimbau。以这两个模型为初始集合,我们扩展了面向葡萄牙语的先进开放编码器生态系统:新增一个15亿参数的高性能驱动大型模型,以及一个1亿参数的高效驱动小型模型。在实现这一核心目标的同时,我们还取得了与生态系统相关的其他重要成果,包括基于SuperGLUE基准的葡萄牙语新数据集——该数据集同样以开放形式发布。