One of the major challenges that under-represented and endangered language communities face in language technology is the lack or paucity of language data. This is also the case of the Southern varieties of the Kurdish and Laki languages for which very limited resources are available with insubstantial progress in tools. To tackle this, we provide a few approaches that rely on the content of local news websites, a local radio station that broadcasts content in Southern Kurdish and fieldwork for Laki. In this paper, we describe some of the challenges of such under-represented languages, particularly in writing and standardization, and also, in retrieving sources of data and retro-digitizing handwritten content to create a corpus for Southern Kurdish and Laki. In addition, we study the task of language identification in light of the other variants of Kurdish and Zaza-Gorani languages.
翻译:代表性不足和濒危语言社群在语言技术领域面临的主要挑战之一,是语言数据的匮乏或稀缺。库尔德语南部方言和拉基语的情况亦然,目前可用的资源极为有限,相关工具的开发进展甚微。为解决这一问题,我们提出几种方法,这些方法依赖于当地新闻网站的内容、一家使用南库尔德语广播的当地广播电台,以及针对拉基语的田野调查。本文描述了此类代表性不足语言所面临的部分挑战,尤其是书写与标准化方面的难题,以及数据源检索和将手写内容逆向数字化以构建南库尔德语和拉基语语料库的过程。此外,我们还结合库尔德语其他变体及扎扎-戈拉尼语,研究了语言识别任务。