India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this work. Given the linguistic diversity, high-quality and accessible Machine Translation (MT) systems are essential in a country like India. Prior to this work, there was (i) no parallel training data spanning all the 22 languages, (ii) no robust benchmarks covering all these languages and containing content relevant to India, and (iii) no existing translation models which support all the 22 scheduled languages of India. In this work, we aim to address this gap by focusing on the missing pieces required for enabling wide, easy, and open access to good machine translation systems for all 22 scheduled Indian languages. We identify four key areas of improvement: curating and creating larger training datasets, creating diverse and high-quality benchmarks, training multilingual models, and releasing models with open access. Our first contribution is the release of the Bharat Parallel Corpus Collection (BPCC), the largest publicly available parallel corpora for Indic languages. BPCC contains a total of 230M bitext pairs, of which a total of 126M were newly added, including 644K manually translated sentence pairs created as part of this work. Our second contribution is the release of the first n-way parallel benchmark covering all 22 Indian languages, featuring diverse domains, Indian-origin content, and source-original test sets. Next, we present IndicTrans2, the first model to support all 22 languages, surpassing existing models on multiple existing and new benchmarks created as a part of this work. Lastly, to promote accessibility and collaboration, we release our models and associated data with permissive licenses at https://github.com/ai4bharat/IndicTrans2.
翻译:印度拥有丰富的语言景观,涵盖四大语系的数十种语言,使用者超过十亿人。其中22种语言被列入印度宪法(称为表列语言),本文聚焦于此。鉴于语言多样性,高质量且易于获取的机器翻译系统在印度这样的国家至关重要。此前,尚存在以下空白:(i) 缺乏覆盖全部22种语言的平行训练数据;(ii) 缺少涵盖所有语言且包含印度本土内容的稳健基准;(iii) 缺乏支持印度全部22种表列语言的翻译模型。本研究旨在填补这些空白,聚焦于实现面向所有22种印度表列语言的广泛、便捷、开放的高质量机器翻译系统所需的关键要素。我们确定了四个改进方向:整理并构建更大规模的训练数据集、创建多样化高质量的基准、训练多语言模型,以及开放访问模型。第一项贡献是发布了Bharat平行语料库,这是目前最大的公开印度语言平行语料库,包含2.3亿个双语句对,其中新增1.26亿个句对,包含本工作创建的64.4万个人工翻译句对。第二项贡献是发布了首个覆盖全部22种印度语言的n路平行基准,涵盖多样领域、印度本土内容及源头原创测试集。第三,我们提出IndicTrans2——首个支持全部22种语言的模型,在多个现有及本工作新创建的基准上超越现有模型。最后,为促进可访问性与协作,我们通过https://github.com/ai4bharat/IndicTrans2以宽松许可证发布模型及相关数据。