Polyglot is a pioneering project aimed at enhancing the non-English language performance of multilingual language models. Despite the availability of various multilingual models such as mBERT (Devlin et al., 2019), XGLM (Lin et al., 2022), and BLOOM (Scao et al., 2022), researchers and developers often resort to building monolingual models in their respective languages due to the dissatisfaction with the current multilingual models non-English language capabilities. Addressing this gap, we seek to develop advanced multilingual language models that offer improved performance in non-English languages. In this paper, we introduce the Polyglot Korean models, which represent a specific focus rather than being multilingual in nature. In collaboration with TUNiB, our team collected 1.2TB of Korean data meticulously curated for our research journey. We made a deliberate decision to prioritize the development of Korean models before venturing into multilingual models. This choice was motivated by multiple factors: firstly, the Korean models facilitated performance comparisons with existing multilingual models; and finally, they catered to the specific needs of Korean companies and researchers. This paper presents our work in developing the Polyglot Korean models, which propose some steps towards addressing the non-English language performance gap in multilingual language models.
翻译:Polyglot是一项旨在提升多语言语言模型非英语语言性能的开创性项目。尽管已有mBERT(Devlin等,2019)、XGLM(Lin等,2022)和BLOOM(Scao等,2022)等多种多语言模型,但研究人员和开发者常因对现有多语言模型非英语语言能力不满,转而构建各自语言的单语模型。为解决这一差距,我们致力于开发性能更优的非英语多语言语言模型。本文介绍的Polyglot韩语模型并非多语言性质,而是聚焦于特定语言。在与TUNiB的合作中,我们团队精心收集了1.2TB韩语数据用于研究。我们刻意优先推进韩语模型开发,而后再涉足多语言模型。这一决策基于多重因素:首先,韩语模型便于与现有模型进行性能对比;其次,它们满足了韩国企业和研究人员的特定需求。本文呈现了我们在开发Polyglot韩语模型方面的工作,这些尝试是弥合多语言语言模型非英语语言性能差距的重要步骤。