We present NusaCrowd, a collaborative initiative to collect and unify existing resources for Indonesian languages, including opening access to previously non-public resources. Through this initiative, we have brought together 137 datasets and 118 standardized data loaders. The quality of the datasets has been assessed manually and automatically, and their value is demonstrated through multiple experiments. NusaCrowd's data collection enables the creation of the first zero-shot benchmarks for natural language understanding and generation in Indonesian and the local languages of Indonesia. Furthermore, NusaCrowd brings the creation of the first multilingual automatic speech recognition benchmark in Indonesian and the local languages of Indonesia. Our work strives to advance natural language processing (NLP) research for languages that are under-represented despite being widely spoken.
翻译:我们提出NusaCrowd协作倡议,旨在收集并整合印度尼西亚语现有资源,包括开放此前非公开资源的访问权限。通过该倡议,我们汇集了137个数据集与118个标准化数据加载器。数据集质量经过人工与自动双重评估,其价值通过多项实验得到验证。NusaCrowd的数据集合使得首次针对印度尼西亚语及当地语言的零样本自然语言理解与生成基准得以创立。此外,该倡议还催生了首个涵盖印度尼西亚语及当地语言的多语言自动语音识别基准。我们的工作致力于推动那些使用广泛但尚未被充分研究的语言在自然语言处理(NLP)领域的研究进展。