While major languages often enjoy substantial attention and resources, the linguistic diversity across the globe encompasses a multitude of smaller, indigenous, and regional languages that lack the same level of computational support. One such region is the Caribbean. While commonly labeled as "English speaking", the ex-British Caribbean region consists of a myriad of Creole languages thriving alongside English. In this paper, we present Guylingo: a comprehensive corpus designed for advancing NLP research in the domain of Creolese (Guyanese English-lexicon Creole), the most widely spoken language in the culturally rich nation of Guyana. We first outline our framework for gathering and digitizing this diverse corpus, inclusive of colloquial expressions, idioms, and regional variations in a low-resource language. We then demonstrate the challenges of training and evaluating NLP models for machine translation in Creole. Lastly, we discuss the unique opportunities presented by recent NLP advancements for accelerating the formal adoption of Creole languages as official languages in the Caribbean.
翻译:尽管主要语言通常获得大量关注和资源,全球语言多样性仍包含众多缺乏同等计算支持的小型、本土及区域性语言。加勒比地区正是这样的区域之一。虽然常被标记为"英语区",前英属加勒比地区实际上并存着大量克里奥尔语与英语。本文提出Guylingo:一个面向克里奥尔语(圭亚那英语词汇克里奥尔语)自然语言处理研究的综合语料库,这种语言是文化丰富的圭亚那使用最广泛的语言。我们首先概述了收集和数字化这一多样化语料库的框架,涵盖低资源语言中的口语表达、习语及地域变体。随后展示了训练和评估克里奥尔语机器翻译NLP模型的挑战。最后,我们探讨了近期NLP进展为加速加勒比地区正式采用克里奥尔语作为官方语言所提供的独特机遇。