Natural language processing (NLP) has a significant impact on society via technologies such as machine translation and search engines. Despite its success, NLP technology is only widely available for high-resource languages such as English and Chinese, while it remains inaccessible to many languages due to the unavailability of data resources and benchmarks. In this work, we focus on developing resources for languages in Indonesia. Despite being the second most linguistically diverse country, most languages in Indonesia are categorized as endangered and some are even extinct. We develop the first-ever parallel resource for 10 low-resource languages in Indonesia. Our resource includes datasets, a multi-task benchmark, and lexicons, as well as a parallel Indonesian-English dataset. We provide extensive analyses and describe the challenges when creating such resources. We hope that our work can spark NLP research on Indonesian and other underrepresented languages.
翻译:自然语言处理(NLP)通过机器翻译和搜索引擎等技术对社会产生了重大影响。尽管取得了成功,但NLP技术目前仅广泛适用于英语和中文等高资源语言,而由于缺乏数据资源和基准测试,许多语言仍无法被覆盖。在本研究中,我们聚焦于为印度尼西亚的语言开发资源。尽管印度尼西亚是语言多样性第二高的国家,但其大多数语言被列为濒危语言,部分甚至已经灭绝。我们首次为印度尼西亚的10种低资源语言构建了平行资源。该资源包含数据集、多任务基准测试、词汇表,以及一个平行印尼语-英语数据集。我们提供了详尽的分析,并描述了创建此类资源时面临的挑战。我们希望本研究能够推动印尼语及其他代表性不足语言的NLP研究。