We present Samanantar, the largest publicly available parallel corpora collection for Indic languages. The collection contains a total of 49.7 million sentence pairs between English and 11 Indic languages (from two language families). Specifically, we compile 12.4 million sentence pairs from existing, publicly-available parallel corpora, and additionally mine 37.4 million sentence pairs from the web, resulting in a 4x increase. We mine the parallel sentences from the web by combining many corpora, tools, and methods: (a) web-crawled monolingual corpora, (b) document OCR for extracting sentences from scanned documents, (c) multilingual representation models for aligning sentences, and (d) approximate nearest neighbor search for searching in a large collection of sentences. Human evaluation of samples from the newly mined corpora validate the high quality of the parallel sentences across 11 languages. Further, we extract 83.4 million sentence pairs between all 55 Indic language pairs from the English-centric parallel corpus using English as the pivot language. We trained multilingual NMT models spanning all these languages on Samanantar, which outperform existing models and baselines on publicly available benchmarks, such as FLORES, establishing the utility of Samanantar. Our data and models are available publicly at https://ai4bharat.iitm.ac.in/samanantar and we hope they will help advance research in NMT and multilingual NLP for Indic languages.
翻译:我们提出Samansantar,这是目前规模最大的公开印度语言平行语料库合集。该合集共包含4970万个英语与11种印度语言(分属两个语系)的句对。具体而言,我们从现有公开平行语料库中编译了1240万个句对,并通过网络挖掘额外获取了3740万个句对,实现总量增长4倍。我们通过整合多种语料库、工具与方法来挖掘网络平行句:(a)网络抓取的单语语料库,(b)用于从扫描文档中提取句子的文档OCR技术,(c)用于句子对齐的多语言表示模型,以及(d)用于大型句子集合搜索的近似最近邻搜索。对新挖掘语料样本的人工评估验证了这11种语言平行句的高质量。此外,我们以英语为枢轴语言,从以英语为中心的平行语料库中提取了所有55种印度语言对之间的8340万个句对。基于Samansantar数据,我们训练了覆盖上述所有语言的多语言神经机器翻译(NMT)模型,该模型在FLORES等公开基准测试中优于现有模型与基线,证实了Samansantar的实用价值。我们的数据与模型已在https://ai4bharat.iitm.ac.in/samanantar 公开提供,期望能推动印度语言NMT与多语言自然语言处理领域的研究进展。