Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic languages by making contributions along 3 important axes (i) monolingual corpora (ii) NLU testsets (iii) multilingual LLMs focusing on Indic languages. Specifically, we curate the largest monolingual corpora, IndicCorp, with 20.9B tokens covering 24 languages from 4 language families - a 2.3x increase over prior work, while supporting 12 additional languages. Next, we create a human-supervised benchmark, IndicXTREME, consisting of nine diverse NLU tasks covering 20 languages. Across languages and tasks, IndicXTREME contains a total of 105 evaluation sets, of which 52 are new contributions to the literature. To the best of our knowledge, this is the first effort towards creating a standard benchmark for Indic languages that aims to test the multilingual zero-shot capabilities of pretrained language models. Finally, we train IndicBERT v2, a state-of-the-art model supporting all the languages. Averaged across languages and tasks, the model achieves an absolute improvement of 2 points over a strong baseline. The data and models are available at https://github.com/AI4Bharat/IndicBERT.
翻译:构建面向印度语言的自然语言理解能力至关重要,这些语言拥有超过十亿的总体使用者。在本文中,我们旨在通过三个关键方向的贡献来提升印度语言的NLU能力:(i)单语语料库(ii)NLU测试集(iii)聚焦印度语言的多语言大语言模型。具体而言,我们整理了最大的单语语料库IndicCorp,包含209亿词元,覆盖4个语系的24种语言——较先前工作增长2.3倍,同时额外支持12种语言。接下来,我们创建了人工监督的基准测试IndicXTREME,涵盖覆盖20种语言的九个多样化NLU任务。跨语言和任务,IndicXTREME共包含105个评估集,其中52个为文献新增贡献。据我们所知,这是首个旨在测试预训练语言模型多语言零样本能力的印度语言标准基准构建工作。最后,我们训练了IndicBERT v2,一个支持所有语言的最先进模型。平均跨语言和任务,该模型在强基线上实现了2个百分点的绝对提升。数据和模型已发布于https://github.com/AI4Bharat/IndicBERT。