We create publicly available language identification (LID) datasets and models in all 22 Indian languages listed in the Indian constitution in both native-script and romanized text. First, we create Bhasha-Abhijnaanam, a language identification test set for native-script as well as romanized text which spans all 22 Indic languages. We also train IndicLID, a language identifier for all the above-mentioned languages in both native and romanized script. For native-script text, it has better language coverage than existing LIDs and is competitive or better than other LIDs. IndicLID is the first LID for romanized text in Indian languages. Two major challenges for romanized text LID are the lack of training data and low-LID performance when languages are similar. We provide simple and effective solutions to these problems. In general, there has been limited work on romanized text in any language, and our findings are relevant to other languages that need romanized language identification. Our models are publicly available at https://github.com/AI4Bharat/IndicLID under open-source licenses. Our training and test sets are also publicly available at https://huggingface.co/datasets/ai4bharat/Bhasha-Abhijnaanam under open-source licenses.
翻译:摘要:我们构建了面向印度宪法所列全部22种印度语言的、公开可用的语言识别(LID)数据集与模型,涵盖原生文字与罗马化文本两种形式。首先,我们创建了Bhasha-Abhijnaanam——一个针对所有22种印度语言的原生文字及罗马化文本的语言识别测试集。其次,我们训练了IndicLID,该语言识别器支持上述所有语言的原生与罗马化文字。对于原生文字文本,IndicLID的语言覆盖范围优于现有LID工具,且性能具有竞争力或更优;而IndicLID是首个面向印度语言罗马化文本的LID工具。罗马化文本LID面临两大主要挑战:训练数据匮乏以及相似语言的低识别性能。我们针对这些问题提出了简单有效的解决方案。总体而言,目前针对任意语言罗马化文本的研究工作十分有限,我们的发现对需要罗马化语言识别的其他语言具有参考价值。我们的模型已基于开源许可协议公开于https://github.com/AI4Bharat/IndicLID。训练集与测试集亦基于开源许可协议发布于https://huggingface.co/datasets/ai4bharat/Bhasha-Abhijnaanam。